Skip to main content

Hugging Face (Inference)

Hugging Face Inference Providers offer OpenAI-compatible chat completions through a single router API. You get access to many models, including DeepSeek, Llama, and more, with one token. Fased uses the OpenAI-compatible endpoint for chat completions only. For text-to-image, embeddings, or speech, use the HF inference clients directly. Hugging Face is a current Fased model provider route. The route id is huggingface; model availability comes from the HF router plus Fased’s built-in catalog fallback.
  • Provider: huggingface
  • Auth: HUGGINGFACE_HUB_TOKEN or HF_TOKEN (fine-grained token with Make calls to Inference Providers)
  • API: OpenAI-compatible (https://router.huggingface.co/v1)
  • Billing: single HF token; pricing follows the active provider route.

Quick start

  1. Create a fine-grained token at Hugging Face → Settings → Tokens with the Make calls to Inference Providers permission.
  2. Run onboarding and choose Hugging Face in the provider dropdown, then enter your API key when prompted:
  1. In onboarding, pick the Hugging Face model you want. The list is loaded from the Inference API when you have a valid token; otherwise a built-in list is shown.
  2. In the browser UI, open Agents, select the Agent, then use Agent > Models to assign Hugging Face as the Agent’s primary, fallback, or task model.
  3. You can also set the Agent default model in config for automation:

Non-interactive example

This stores the key and can set huggingface/openai/gpt-oss-120b as the Agent default model.

Environment note

If the Gateway runs as a daemon (launchd/systemd), make sure HUGGINGFACE_HUB_TOKEN or HF_TOKEN is available to that process. Use ~/.fased/.env or env.shellEnv.

Model discovery and onboarding dropdown

Fased discovers models by calling the Inference endpoint directly:
For the full list, send Authorization: Bearer $HUGGINGFACE_HUB_TOKEN or Authorization: Bearer $HF_TOKEN. Some endpoints return a subset without auth. The response is OpenAI-style:
When you configure a Hugging Face API key through onboarding, HUGGINGFACE_HUB_TOKEN, or HF_TOKEN, Fased uses this request to discover available chat-completion models. During interactive onboarding, after you enter your token, the Default Hugging Face model dropdown is populated from that list. If the request fails, Fased uses the built-in catalog. At runtime, for example during Gateway startup, Fased calls GET https://router.huggingface.co/v1/models again when a key is present. The result is merged with the built-in catalog for metadata such as context window and cost. If the request fails or no key is set, only the built-in catalog is used.

Model names and editable options

  • Name from API: The model display name is hydrated from GET /v1/models when the API returns name, title, or display_name. Otherwise it is derived from the model id, for example openai/gpt-oss-120b → “GPT OSS 120B”.
  • Override display name: You can set a custom label per model in config so it appears the way you want in the CLI and UI:
  • Provider / policy selection: Append a suffix to the model id to choose how the router picks the backend:
    • :fastest — highest throughput (router picks; provider choice is locked — no interactive backend picker).
    • :cheapest — lowest cost per output token (router picks; provider choice is locked).
    • :provider — force a specific backend (e.g. :sambanova, :together).
    When you select :cheapest or :fastest, the provider is locked: the router decides by cost or speed and no optional “prefer specific backend” step is shown. You can add these as separate entries in models.providers.huggingface.models or set model.primary with the suffix. You can also set your default order in Inference Provider settings. No suffix means Fased uses that order.
  • Config merge: Existing entries in models.providers.huggingface.models, for example in models.json, are kept when config is merged. Custom name, alias, and model options stay in place.

Model IDs and configuration examples

Model refs use the form huggingface/<org>/<model> with Hub-style IDs. The first-run list is curated from the official Chat Completion recommendations. When you have a valid token, the runtime can still discover more with GET https://router.huggingface.co/v1/models. Example IDs from Fased’s built-in Hugging Face catalog: You can append :fastest, :cheapest, or :provider such as :together or :sambanova to the model id. Set your default order in Inference Provider settings. See Inference Providers and GET https://router.huggingface.co/v1/models for the full list.

Complete configuration examples

Primary GPT-OSS 120B with Qwen Coder fallback:
Qwen as default, with :cheapest and :fastest variants:
GPT-OSS + GLM + DeepSeek with aliases:
Force a specific backend with :provider:
Multiple Qwen and DeepSeek models with policy suffixes: