Hugging Face (Inference)
Hugging Face Inference Providers offer OpenAI-compatible chat completions through a single router API. You get access to many models, including DeepSeek, Llama, and more, with one token. Fased uses the OpenAI-compatible endpoint for chat completions only. For text-to-image, embeddings, or speech, use the HF inference clients directly. Hugging Face is a current Fased model provider route. The route id ishuggingface; model availability comes from the HF router plus Fased’s built-in
catalog fallback.
- Provider:
huggingface - Auth:
HUGGINGFACE_HUB_TOKENorHF_TOKEN(fine-grained token with Make calls to Inference Providers) - API: OpenAI-compatible (
https://router.huggingface.co/v1) - Billing: single HF token; pricing follows the active provider route.
Quick start
- Create a fine-grained token at Hugging Face → Settings → Tokens with the Make calls to Inference Providers permission.
- Run onboarding and choose Hugging Face in the provider dropdown, then enter your API key when prompted:
- In onboarding, pick the Hugging Face model you want. The list is loaded from the Inference API when you have a valid token; otherwise a built-in list is shown.
- In the browser UI, open Agents, select the Agent, then use Agent > Models to assign Hugging Face as the Agent’s primary, fallback, or task model.
- You can also set the Agent default model in config for automation:
Non-interactive example
huggingface/openai/gpt-oss-120b as the Agent
default model.
Environment note
If the Gateway runs as a daemon (launchd/systemd), make sureHUGGINGFACE_HUB_TOKEN or HF_TOKEN is available to that process. Use
~/.fased/.env or env.shellEnv.
Model discovery and onboarding dropdown
Fased discovers models by calling the Inference endpoint directly:Authorization: Bearer $HUGGINGFACE_HUB_TOKEN or
Authorization: Bearer $HF_TOKEN. Some endpoints return a subset without auth.
The response is OpenAI-style:
HUGGINGFACE_HUB_TOKEN, or HF_TOKEN, Fased uses this request to discover
available chat-completion models.
During interactive onboarding, after you enter your token, the Default
Hugging Face model dropdown is populated from that list. If the request fails,
Fased uses the built-in catalog.
At runtime, for example during Gateway startup, Fased calls
GET https://router.huggingface.co/v1/models again when a key is present. The
result is merged with the built-in catalog for metadata such as context window
and cost. If the request fails or no key is set, only the built-in catalog is
used.
Model names and editable options
- Name from API: The model display name is hydrated from
GET /v1/modelswhen the API returnsname,title, ordisplay_name. Otherwise it is derived from the model id, for exampleopenai/gpt-oss-120b→ “GPT OSS 120B”. - Override display name: You can set a custom label per model in config so it appears the way you want in the CLI and UI:
-
Provider / policy selection: Append a suffix to the model id to choose
how the router picks the backend:
:fastest— highest throughput (router picks; provider choice is locked — no interactive backend picker).:cheapest— lowest cost per output token (router picks; provider choice is locked).:provider— force a specific backend (e.g.:sambanova,:together).
models.providers.huggingface.modelsor setmodel.primarywith the suffix. You can also set your default order in Inference Provider settings. No suffix means Fased uses that order. -
Config merge: Existing entries in
models.providers.huggingface.models, for example inmodels.json, are kept when config is merged. Customname,alias, and model options stay in place.
Model IDs and configuration examples
Model refs use the formhuggingface/<org>/<model> with Hub-style IDs. The
first-run list is curated from the official Chat Completion recommendations.
When you have a valid token, the runtime can still discover more with
GET https://router.huggingface.co/v1/models.
Example IDs from Fased’s built-in Hugging Face catalog:
You can append
:fastest, :cheapest, or :provider such as :together or
:sambanova to the model id. Set your default order in
Inference Provider settings.
See Inference Providers and
GET https://router.huggingface.co/v1/models for the full list.