Skip to content
Back to skills

Nodetool Model Provider Config

CSecurity

Configure NodeTool model providers, credentials, and model selection, including local Ollama and Hugging Face models.

  • 552 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 1, 2026
ai-agentstypescriptgobashnodeapibackendperformance

Works with

  • cli
  • api

Security analysis

C71/100
  • criticalPipes output to a shell interpreter
  • mediumUses curl or wget to download content
  • criticalDownloads and executes remote scripts — classic supply chain attack

Pro shows the line behind each finding and how to fix it

Scanned October 6, 2026

npx -y skills add nodetool-ai/nodetool --skill nodetool-model-provider-config --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Nodetool Model Provider Config?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Nodetool Model Provider Config
[![Security: C — Skills Directory](https://www.skillsdirectory.com/api/skills/nodetool-ai-nodetool-model-provider-config/badge)](https://www.skillsdirectory.com/skills/nodetool-ai-nodetool-model-provider-config)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: nodetool-model-provider-config
description: "Configure NodeTool model providers, credentials, and model selection, including local Ollama and Hugging Face models."
featured: true
---

You help users configure AI model providers and select the right models for their tasks.

# Provider Overview

| Provider | Type | Key Env Var | Models |
|----------|------|-------------|--------|
| **OpenAI** | Cloud | `OPENAI_API_KEY` | GPT-5.4 / GPT-5.4-mini, GPT-Image, TTS, Whisper |
| **Anthropic** | Cloud | `ANTHROPIC_API_KEY` | Claude Sonnet 4.6, Haiku |
| **Gemini (Google)** | Cloud | `GEMINI_API_KEY` | Gemini 2.5, Veo, Nano Banana |
| **xAI** | Cloud | `XAI_API_KEY` | Grok 4 |
| **Ollama** | Local | `OLLAMA_API_URL` | Qwen, Llama 3, Mistral, any GGUF |
| **HuggingFace** | Local/Cloud | `HF_TOKEN` | 1000+ models, auto-download |
| **FAL** | Cloud | `FAL_API_KEY` | Fast image/video generation |
| **Replicate** | Cloud | `REPLICATE_API_TOKEN` | Community models |
| **Comfy Cloud** | Cloud | `COMFY_API_KEY` | No model list; runs a ComfyUI workflow on Comfy's GPUs via `lib.comfy.RunWorkflowOnCloud` |
| **vLLM** | Local | `VLLM_API_URL` | Self-hosted, OpenAI-compatible |
| **llama.cpp** | Local | — | GGUF models, CPU/GPU |
| **MLX** | Local | — | Apple Silicon optimized |

Other registered chat providers (any of these is valid for `-p/--provider`):
`groq`, `mistral`, `deepseek`, `moonshot`, `minimax`, `cerebras`, `alibaba`
(`DASHSCOPE_API_KEY`, Qwen via Alibaba Cloud Model Studio), `together`,
`openrouter`, `requesty`, `codex`, `claude_agent_sdk`, `lmstudio`. Run `nodetool models
providers` to see configured providers and `nodetool models recommended` for the
curated model list.

# API Key Setup

```bash
# Via CLI (encrypted storage)
nodetool secrets store OPENAI_API_KEY
nodetool secrets store ANTHROPIC_API_KEY
nodetool secrets store GEMINI_API_KEY
nodetool secrets store HF_TOKEN
nodetool secrets store FAL_API_KEY
nodetool secrets store REPLICATE_API_TOKEN
nodetool secrets store COMFY_API_KEY

# Via environment variables
export OPENAI_API_KEY=sk-...
export ANTHROPIC_API_KEY=sk-ant-...
export GEMINI_API_KEY=AI...
export HF_TOKEN=hf_...
export FAL_API_KEY=...
export REPLICATE_API_TOKEN=r8_...
export COMFY_API_KEY=comfyui-...
export OLLAMA_API_URL=http://localhost:11434
```

# Model Selection by Task

## whisper.cpp speech recognition

Both providers are local-only and unavailable on the cloud profile.
`whisper_cpp` runs speech recognition in the backend through the optional
`@fugood/whisper.node@1.1.3` runtime package. Install **whisper.cpp** from the
Package Manager, then download a GGML model from Models. It serves
`nodetool.text.AutomaticSpeechRecognition` and `whisper_cpp.LiveTranscription`.
Use the absolute model path returned by discovery unchanged as the ASR model id.

Models use the Hugging Face hub cache (`HF_HUB_CACHE`, then `HF_HOME/hub`, then
`~/.cache/huggingface/hub`). `WHISPER_CPP_MODELS_DIR` adds another directory.
`WHISPER_CPP_GPU_BACKEND` accepts `auto`, `metal`, `cuda`, `vulkan`, or `cpu`.
`auto` uses the default build, including Metal on macOS. Restart the backend
after changing the backend setting. Live transcription accepts base64 PCM16
mono chunks with an optional `content_metadata.sample_rate` and uses Silero
VAD when installed, otherwise fixed windows.

`whisper_cpp_server` calls a user-run `whisper-server`. Set
`WHISPER_CPP_SERVER_URL` to its base URL. The model id is `default`, representing
the model loaded by that server. It requires no native runtime package.

## Language / Chat

| Need | Model | Provider | Notes |
|------|-------|----------|-------|
| Best quality | gpt-5.4, claude-sonnet-4-6 | OpenAI, Anthropic | Highest capability |
| Good balance | gpt-5.4-mini, gemini-2.5-flash | OpenAI, Gemini | Fast + cheap |
| Local/private | Llama 3.3 70B, Qwen 3.5 | Ollama | No data leaves machine |
| Lightweight local | Llama 3 8B, Mistral 7B | Ollama | Low memory |
| Code | claude-sonnet-4-6, gpt-5.4 | Anthropic, OpenAI | Best for coding |

## Image Generation

| Need | Model | Provider | Notes |
|------|-------|----------|-------|
| Best quality | FLUX.2 Dev | HuggingFace, FAL | State-of-art |
| Fast | FLUX Schnell | HuggingFace | Quick iterations |
| Versatile | SDXL | HuggingFace | Many LoRAs available |
| API-based | GPT Image 2, Nano Banana | OpenAI, Gemini/KIE | No local GPU needed |

## Video Generation

| Need | Model | Provider |
|------|-------|----------|
| Best quality | Sora 2 Pro | OpenAI (KIE) |
| Fast | Wan 2.6 | KIE |
| Image-to-video | Kling 2.6 | KIE |
| Talking avatar | Kling AI Avatar | KIE |

## Speech & Audio

| Need | Model | Provider |
|------|-------|----------|
| TTS (quality) | ElevenLabs | ElevenLabs |
| TTS (fast/free) | Whisper TTS | HuggingFace |
| ASR (accuracy) | Whisper Large V3 | HuggingFace |
| ASR (fast) | Whisper Turbo | HuggingFace |

## Embeddings

| Need | Model | Provider |
|------|-------|----------|
| General text | text-embedding-3-small | OpenAI |
| Best quality | text-embedding-3-large | OpenAI |
| Local/free | sentence-transformers | HuggingFace |

# Local Model Setup

## Ollama (Easiest)

```bash
# Install
curl -fsSL https://ollama.ai/install.sh | sh

# Pull models
ollama pull llama3
ollama pull mistral
ollama pull qwen2

# Verify
ollama list

# NodeTool auto-discovers Ollama at localhost:11434
# Override: export OLLAMA_API_URL=http://host:11434
```

## HuggingFace (Auto-Download)

Models auto-download to `~/.cache/huggingface/` on first use.

For gated models:
1. Accept terms on HuggingFace Hub
2. Set `HF_TOKEN`
3. Model downloads automatically

## llama.cpp

Manual GGUF model loading. Best for CPU inference and quantized models.

## MLX (Apple Silicon only)

Optimized for M1/M2/M3 chips. Lower memory usage than standard PyTorch.

# Local Inference Performance

| Framework | Throughput | Memory | Hardware |
|-----------|-----------|--------|----------|
| **llama.cpp** | Medium | Excellent | CPU, GPU |
| **MLX** | Good | Excellent | Apple Silicon |
| **Nunchaku** | Excellent | Excellent | NVIDIA GPU |
| **Transformers** | Medium | Good | Any |

# Provider-Agnostic Nodes

These nodes work with any provider — just select the model:

| Node | Purpose |
|------|---------|
| `nodetool.agents.Agent` | Any LLM for chat/reasoning |
| `nodetool.image.TextToImage` | Any image generation model |
| `nodetool.image.ImageToImage` | Any image transformation model |
| `nodetool.video.TextToVideo` | Any video generation model |
| `nodetool.video.ImageToVideo` | Any image-to-video model |
| `nodetool.audio.TextToSpeech` | Any TTS model |
| `nodetool.text.AutomaticSpeechRecognition` | Any ASR model |

# Custom Provider Development

Providers extend `BaseProvider` from `@nodetool-ai/runtime` (not `@nodetool-ai/core`).
Both `generateMessage` and `generateMessages` take a single **args object**.

```typescript
import {
  BaseProvider,
  type ProviderId,
  type Message,
  type ProviderStreamItem,
  type ProviderTool,
  type LanguageModel,
} from "@nodetool-ai/runtime";

export class MyProvider extends BaseProvider {
  private apiKey: string;

  constructor(kwargs: Record<string, unknown> = {}) {
    super("my_provider" as ProviderId);
    this.apiKey = String(kwargs["MY_API_KEY"] ?? process.env.MY_API_KEY ?? "");
  }

  static override requiredSecrets(): string[] {
    return ["MY_API_KEY"];
  }

  // Non-streaming: return a single assistant Message.
  async generateMessage(args: {
    messages: Message[];
    model: string;
    tools?: ProviderTool[];
  }): Promise<Message> {
    // Call your API with args.messages / args.model …
    return { role: "assistant", content: "response text" };
  }

  // Streaming: yield ProviderStreamItem chunks.
  async *generateMessages(args: {
    messages: Message[];
    model: string;
    tools?: ProviderTool[];
  }): AsyncGenerator<ProviderStreamItem> {
    yield { type: "chunk", content: "response text", done: false };
  }

  override async getAvailableLanguageModels(): Promise<LanguageModel[]> {
    return [{ id: "my-model", name: "My Model", provider: "my_provider" }];
  }
}
```

Register it with `registerProvider("my_provider", MyProvider)` from
`@nodetool-ai/runtime`.

# Provider Capabilities

| Capability | OpenAI | Anthropic | Google | Ollama | HF |
|-----------|--------|-----------|--------|--------|-----|
| Chat/Text | yes | yes | yes | yes | yes |
| Vision | yes | yes | yes | some | yes |
| Image Gen | yes (GPT-Image) | no | no | no | yes |
| Video Gen | no | no | yes (Veo) | no | some |
| TTS | yes | no | no | no | yes |
| ASR | yes (Whisper) | no | no | no | yes |
| Embeddings | yes | no | yes | yes | yes |
| Tool Calling | yes | yes | yes | some | no |

# Common Pitfalls

- **Wrong key env var name**: Each provider has a specific name (see table above)
- **Ollama not running**: Start with `ollama serve` before using
- **Gated HF models**: Must accept terms on hub.huggingface.co first
- **GPU memory**: Large models need 8-24GB VRAM; use quantized versions
- **Rate limits**: Cloud providers have rate limits; implement retries or use local
- **Model ID mismatch**: Use the exact model ID from the provider (e.g., `gpt-5.4`, `claude-sonnet-4-6`) — `nodetool models by-provider <provider>` lists them

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…