Hugging Face
Use the hf provider for Hugging Face Inference Providers.
Use hf.<model_name>[:provider] to specify models. If no provider suffix is supplied, Hugging Face auto-routes the request.
fast-agent --model kimi
fast-agent --model kimi26instant
fast-agent --model hf.openai/gpt-oss-120b
fast-agent --model hf.moonshotai/kimi-k2-instruct-0905:groq
fast-agent --model "hf.moonshotai/Kimi-K2.6:novita?reasoning=on"
Curated aliases such as kimi, deepseek-hf, glm, and minimax include provider choices and request defaults tested with fast-agent features such as structured outputs and tool use. Capability can still vary by backing provider.
Qwen3.8 27B
Use the canonical Qwen3.8 27B model ID with the Hugging Face router:
fast-agent --model "hf.Qwen/Qwen3.8-27B?reasoning=xhigh"
fast-agent --model "hf.Qwen/Qwen3.8-27B?reasoning=low"
fast-agent --model "hf.Qwen/Qwen3.8-27B?reasoning=off"
The model has a native 262,144-token context window. Its optional 1M context
requires an explicitly configured YaRN deployment; dedicated endpoints should
use the limit returned by their own /v1/models response.
Qwen3.8 thinking is enabled by default. Fast-agent supports the documented
low, medium, and xhigh reasoning efforts. Although the upstream model
defaults to xhigh, fast-agent uses medium: repeated fixed-answer trials
matched xhigh correctness with materially lower completion usage and latency.
Use reasoning=xhigh when maximum deliberation matters more than throughput.
reasoning=off sends chat_template_kwargs.enable_thinking: false. Historical
reasoning is replayed as reasoning_content, preserving the model's default
multi-turn thinking behavior.
Qwen3.8 uses the grok_shell execution profile by default. It exposes shell
with explicit timeout, background, and working-directory fields alongside the
unified process tool. Set shell_execution.tool_profile explicitly to
override automatic model selection.
The model card recommends temperature=1.0, top_p=0.95, top_k=20,
min_p=0.0, presence_penalty=0.0, and repetition_penalty=1.0 for thinking
mode. Put these values in a model overlay when a dedicated endpoint should use
them by default.
The model profile supports text, JPEG/PNG/WebP images, and MP4 video. Live fast-agent onboarding on August 22, 2026 verified streamed reasoning, reasoning disable, shell tool execution and continuation, JSON-object structured output, tool-assisted structured output, local image and video attachments, and usage accounting on a dedicated Hugging Face OpenAI Chat Completions endpoint. JSON object mode is advertised; JSON Schema mode is not.
For tool-assisted structured output, fast-agent defers JSON-object enforcement until after the tool result. In repeated live trials, enforcing JSON while tools were still active caused the model to fabricate an answer instead of calling the tool; the deferred workflow called the tool and preserved its payload reliably. This policy performs one tool-gathering phase followed by a schema-only final phase; workflows that require sequential dependent tool rounds should gather those results before requesting the structured final.
Muse Glimmer via Together
glimmer routes Muse Glimmer 30B
through the Hugging Face Inference Providers router using Together:
The preset resolves to hf.meta-models/Muse-Glimmer-30B:together and applies
Meta's recommended sampling defaults: temperature=1.0, top_p=0.95, and
top_k=64.
Muse Glimmer supports text and image input with text output and a 131,072-token
context window. Its reasoning control is a chat-template setting rather than an
OpenAI reasoning_effort field. Fast-agent maps low, medium, high, and
xhigh to chat_template_kwargs.reasoning_strength; the default is high.
Meta's released chat template and Together's model page describe tool calling, while Together's serverless model catalog currently marks function calling and structured outputs as unavailable for this endpoint. Live fast-agent testing confirms regular streamed tool calls and tool-result continuation work; the Hugging Face adapter uses manual stream accumulation for Together's null-valued tool-call continuation fragments.
Fast-agent does not advertise structured JSON support for glimmer. Prompted
JSON and tool-assisted JSON can succeed, but native Pydantic structured output
was not reliable in live testing.
Kimi instant mode
Kimi models that support instant mode can disable reasoning with the instant query parameter:
fast-agent --model "hf.moonshotai/Kimi-K2.5?instant=on" # thinking disabled
fast-agent --model "hf.moonshotai/Kimi-K2.5?instant=off" # thinking enabled
Gemma thinking mode
gemma4 routes Gemma 4 31B through Hugging Face Inference Providers on Cerebras:
fast-agent --model gemma4
fast-agent --model "hf.google/gemma-4-31B-it:cerebras?temperature=1.0&top_p=0.95"
Gemma 4 reasoning is disabled by default on Cerebras. Enable it with reasoning_effort
values through fast-agent's reasoning query parameter:
fast-agent --model "gemma4?reasoning=medium" # sends reasoning_effort=medium
fast-agent --model "gemma4?reasoning=none" # sends reasoning_effort=none
Hugging Face MCP authentication
HF_TOKEN is automatically applied when connecting to Hugging Face MCP servers:
hf.co/huggingface.cousesAuthorization: Bearer {HF_TOKEN}*.hf.spaceuses bothAuthorization: Bearer {HF_TOKEN}andX-HF-Authorization: Bearer {HF_TOKEN}
Model aliases
| Model Alias | Maps to |
|---|---|
DeepSeek V4 Flash 0731 (baseten) |
hf.deepseek-ai/DeepSeek-V4-Flash-0731:baseten?max_tokens=384000 |
DeepSeek V4 Flash 0731 (deepinfra) |
hf.deepseek-ai/DeepSeek-V4-Flash-0731:deepinfra |
deepseek-ai/deepseek-v3.1 |
deepseek-ai/deepseek-v3.1 |
deepseek-ai/deepseek-v3.2 |
deepseek-ai/deepseek-v3.2 |
deepseek-ai/deepseek-v4-flash-0731 |
deepseek-ai/deepseek-v4-flash-0731 |
deepseek-ai/deepseek-v4-pro |
deepseek-ai/deepseek-v4-pro |
deepseek-hf |
hf.deepseek-ai/DeepSeek-V4-Pro:together |
deepseek32 |
hf.deepseek-ai/DeepSeek-V3.2:fireworks-ai |
deepseek4-hf |
hf.deepseek-ai/DeepSeek-V4-Pro:together |
deepseek4pro-hf |
hf.deepseek-ai/DeepSeek-V4-Pro:together |
deepseekv4pro-hf |
hf.deepseek-ai/DeepSeek-V4-Pro:together |
gemma4 |
hf.google/gemma-4-31B-it:cerebras?temperature=1.0&top_p=0.95 |
glimmer |
hf.meta-models/Muse-Glimmer-30B:together?temperature=1.0&top_p=0.95&top_k=64 |
glm |
hf.zai-org/GLM-5.2:zai-org |
GLM 5.2 (deepinfra) |
hf.zai-org/GLM-5.2:deepinfra |
GLM 5.2 (fireworks-ai) |
hf.zai-org/GLM-5.2:fireworks-ai |
GLM 5.2 (zai-org) |
hf.zai-org/GLM-5.2:zai-org |
glm47 |
hf.zai-org/GLM-4.7:cerebras |
glm5 |
hf.zai-org/GLM-5:novita |
glm51 |
hf.zai-org/GLM-5.1:together |
glm52 |
hf.zai-org/GLM-5.2:zai-org |
google/gemma-4-31b-it |
google/gemma-4-31b-it |
gpt-oss |
hf.openai/gpt-oss-120b:cerebras |
gpt-oss-20b |
hf.openai/gpt-oss-20b |
kimi |
hf.moonshotai/Kimi-K2.7-Code:fireworks-ai?temperature=1.0&top_p=0.95&reasoning=on |
Kimi K3 (fireworks-ai) |
hf.moonshotai/Kimi-K3:fireworks-ai |
Kimi K3 (together) |
hf.moonshotai/Kimi-K3:together |
kimi-2.5 |
hf.moonshotai/Kimi-K2.5:novita?temperature=1.0&top_p=0.95&reasoning=on |
kimi-2.6 |
hf.moonshotai/Kimi-K2.6:novita?temperature=1.0&top_p=0.95&reasoning=on |
kimi25 |
hf.moonshotai/Kimi-K2.5:novita?temperature=1.0&top_p=0.95&reasoning=on |
kimi25instant |
hf.moonshotai/Kimi-K2.5:novita?temperature=0.6&top_p=0.95&reasoning=off |
kimi26 |
hf.moonshotai/Kimi-K2.6:novita?temperature=1.0&top_p=0.95&reasoning=on |
kimi26instant |
hf.moonshotai/Kimi-K2.6:novita?temperature=0.6&top_p=0.95&reasoning=off |
kimi27 |
hf.moonshotai/Kimi-K2.7-Code:fireworks-ai?temperature=1.0&top_p=0.95&reasoning=on |
kimi27code |
hf.moonshotai/Kimi-K2.7-Code:fireworks-ai?temperature=1.0&top_p=0.95&reasoning=on |
kimithink |
hf.moonshotai/Kimi-K2.6:novita?temperature=1.0&top_p=0.95&reasoning=on |
meta-models/muse-glimmer-30b |
meta-models/muse-glimmer-30b |
minimax |
hf.MiniMaxAI/MiniMax-M3:together?temperature=1.0&top_p=0.95&top_k=40 |
minimax2.5 |
hf.MiniMaxAI/MiniMax-M2.5:novita?temperature=1.0&top_p=0.95&top_k=40 |
minimax21 |
hf.MiniMaxAI/MiniMax-M2.1:novita |
minimax25 |
hf.MiniMaxAI/MiniMax-M2.5:fireworks-ai?temperature=1.0&top_p=0.95&top_k=40 |
minimax27 |
hf.MiniMaxAI/MiniMax-M2.7:fireworks-ai?temperature=1.0&top_p=0.95&top_k=40 |
minimax3 |
hf.MiniMaxAI/MiniMax-M3:together?temperature=1.0&top_p=0.95&top_k=40 |
moonshotai/kimi-k2 |
moonshotai/kimi-k2 |
moonshotai/kimi-k2-instruct-0905 |
moonshotai/kimi-k2-instruct-0905 |
moonshotai/kimi-k2-thinking |
moonshotai/kimi-k2-thinking |
moonshotai/kimi-k2.5 |
moonshotai/kimi-k2.5 |
moonshotai/kimi-k2.6 |
moonshotai/kimi-k2.6 |
moonshotai/kimi-k2.7-code |
moonshotai/kimi-k2.7-code |
moonshotai/kimi-k3 |
moonshotai/kimi-k3 |
qwen/qwen3.5-397b-a17b |
qwen/qwen3.5-397b-a17b |
qwen/qwen3.6-35b-a3b |
qwen/qwen3.6-35b-a3b |
qwen/qwen3.8-27b |
qwen/qwen3.8-27b |
qwen35 |
hf.Qwen/Qwen3.5-397B-A17B:novita?temperature=0.6&top_p=0.95&top_k=20&min_p=0.0&presence_penalty=0.0&repetition_penalty=1.0&reasoning=on |
qwen35instruct |
hf.Qwen/Qwen3.5-397B-A17B:novita?temperature=0.7&top_p=0.8&top_k=20&min_p=0.0&presence_penalty=1.5&repetition_penalty=1.0&reasoning=off |
qwen36 |
hf.Qwen/Qwen3.6-35B-A3B:deepinfra?temperature=0.6&top_p=0.95&top_k=20&min_p=0.0&presence_penalty=0.0&repetition_penalty=1.0&reasoning=on |
qwen36instruct |
hf.Qwen/Qwen3.6-35B-A3B:deepinfra?temperature=0.7&top_p=0.8&top_k=20&min_p=0.0&presence_penalty=1.5&repetition_penalty=1.0&reasoning=off |
zai-org/glm-5.2 |
zai-org/glm-5.2 |