Token Gear by FreshToken · LLM token aggregation & redistribution

One gear for every Token Engine.

Token Gear aggregates model capacity from many NCP Token Engines and redistributes it to tenants through one OpenAI-compatible API — with session-aware routing, SLA and security policy, quotas, and prepaid metering to the token.

API base URL https://api.freshtoken.ai/v1

Why Token Gear

The layer between your tenants and every engine

Token Gear owns business-level identity, routing and policy, so tenants get one reliable API and engines get steady, well-shaped demand.

Unified API

One OpenAI-compatible endpoint for /v1/chat/completions, /v1/completions and /v1/embeddings, streamed over SSE. Keep your SDK; change the base URL.

Session-aware routing & SLA

cheapest, fastest, balanced or pinned, on live health with circuit breakers and failover. Each session stays sticky to one Token Engine.

Tenant quotas & security policy

Organizations as tenants, with per-key quotas, spend caps, rate limits and model allow-lists. Keys are shown once and stored as hashes.

Prepaid metering & settlement

The most a request can cost is reserved before it goes upstream and settled to the token when it ends. Shared org balances, monthly invoices, reseller referrals.

Architecture

From tenant sessions to Token Engines

Token Gear sits between tenants and NCP domains. It keeps the logical session; each Token Engine keeps the execution session. A bidirectional API joins the two.

Individual users

  • Users of Tenant 1
  • Users of Tenant 2
  • Users of Tenant 3

Tenants · business users

  • Tenant 1User sessions
  • Tenant 2User sessions
  • Tenant 3User sessions

Token Gear

Aggregator / Redistribution Platform · session identity, routing, SLA and security policy

Policy & Session Router

  • Routing
  • SLA
  • Security Policy

Logical Session State

  • Tenant identity
  • Session tracking
  • Quotas
API (bidirectional)
  • NCP A

    Token Engine

    Execution Session State
    Inference Scheduler
    • P/D disaggregation
    • KV Cache
    • Batching
    • DP/EP
    LLM hosting
    • DeepSeek
    • Qwen
    GPU / compute
  • NCP B

    Token Engine

    Execution Session State
    Inference Scheduler
    • P/D disaggregation
    • KV Cache
    • Batching
    • DP/EP
    LLM hosting
    • Kimi
    • Qwen
    GPU / compute
  • NCP C

    Token Engine

    Execution Session State
    Inference Scheduler
    • P/D disaggregation
    • KV Cache
    • Batching
    • DP/EP
    LLM hosting
    • Llama
    • DeepSeek
    GPU / compute
Diagram: users of three tenants reach their tenant; tenants send user sessions to Token Gear, which routes them over a bidirectional API to Token Engines in NCP domains A, B and C. Tenant 1 · NCP A Tenant 2 · NCP B · Token Gear Tenant 3 · NCP C Request flow Bidirectional API

Five principles

How capacity is shared

Sharing is what makes redistribution efficient. These five rules decide where it goes, and where it stops.

  1. 01

    Balanced Sharing

    Sharing is optimized across security, efficiency and flexibility — not maximized indiscriminately.

  2. 02

    Spatial Sharing

    One Token Engine can serve sessions from multiple tenants at the same time.

  3. 03

    Temporal Sharing

    Sessions are allocated over time; capacity is reallocated to the next session when one completes.

  4. 04

    Dual-layer Sessions

    Token Gear manages logical sessions at the tenant level; the Token Engine manages execution sessions at the engine level.

  5. 05

    Session Persistence

    Each session is sticky to one Token Engine for its lifetime. No cross-engine migration within a session.

Featured models

Open and frontier models, one key

Open-weight models served by NCP Token Engines, alongside frontier models from provider APIs.

View all in the console →
  • DeepSeek

    deepseek/deepseek-v3

    Context
    128K
    Served by
    NCP ANCP C
  • Moonshot AI

    moonshot/kimi-k2

    Context
    128K
    Served by
    NCP B
  • Qwen

    qwen/qwen3

    Context
    128K
    Served by
    NCP ANCP B
  • Meta

    meta/llama-3.3-70b

    Context
    128K
    Served by
    NCP C
  • OpenAI

    openai/gpt-4o

    Context
    128K
    Served by
    Provider API
  • Anthropic

    anthropic/claude-sonnet-4.5

    Context
    200K
    Served by
    Provider API

Two sides, one platform

Buy tokens, or supply them

Tenants and developers buy tokens through one API. NCPs running Token Engines supply capacity and receive redistributed demand.

For builders & tenants

One key for every engine

  • One OpenAI-compatible API key, one prepaid balance
  • Organizations as tenants, with shared balances and members
  • Per-key quotas, spend caps, rate limits and model allow-lists
  • Usage to the token, and monthly invoices
Get API key

For NCPs / Token Engine operators

Plug your engine into Token Gear

  • Connect through the bidirectional API between Token Gear and your engine
  • Receive redistributed tenant demand, already authenticated and within quota
  • Execution sessions, scheduling and GPUs stay yours
  • Sessions stay sticky to your engine for their lifetime
Talk to us about joining

Get started

Your first request in three steps

Every public endpoint lives under /v1.

  1. Sign up

    Create an account, personal or for your organization.

  2. Buy credits

    Top up a prepaid balance; each request is metered to the token.

  3. Get your API key

    Create a key starting with ft-. It is shown once.

POST /v1/chat/completions
curl https://api.freshtoken.ai/v1/chat/completions \
  -H "Authorization: Bearer ft-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v3",
    "messages": [{"role": "user", "content": "Hello, Token Gear"}],
    "stream": true
  }'

FAQ

Questions, answered

Something else? Write to support@freshtoken.ai.

What is Token Gear?

FreshToken's LLM token aggregation and redistribution platform. It aggregates capacity from many Token Engines and provider APIs, and redistributes it to tenants through one OpenAI-compatible API, handling session identity, routing, SLA and security policy along the way.

How is it different from calling model providers directly?

You hold one key and one balance instead of an account per provider. Token Gear routes each request on live health with failover, applies your tenant's quotas and policies, and meters every request to the token against a single prepaid balance.

What is an NCP, and what is a Token Engine?

An NCP is a compute provider domain. Each NCP runs a Token Engine: execution session state and an inference scheduler (P/D disaggregation, KV cache, batching, DP/EP) on top of its LLM hosting and GPU infrastructure.

How do sessions and stickiness work?

Sessions have two layers. Token Gear keeps the logical session — tenant identity, tracking and quotas. The Token Engine keeps the execution session. A session stays on one Token Engine for its lifetime and is never migrated between engines midway.

How does billing work?

Prepaid credits, metered per token. The most a request can cost is reserved before it goes upstream, and the actual tokens are settled when it ends. Organizations share a balance, and statements are issued monthly.

What happens to my data?

Your content is passed to the engine or provider that serves the request and back to you. We keep what is needed to run and bill an account — keys, usage records and balances — not your prompts or responses. See the privacy policy.

Is it compatible with the OpenAI SDKs?

Yes. Chat completions, completions and embeddings follow the OpenAI format, including SSE streaming. Point the SDK's base URL at Token Gear and use an ft- key.

How does an NCP join?

Write to support@freshtoken.ai. We connect your Token Engine through the bidirectional API, register the models it hosts, and start redistributing tenant demand to it.

Put every Token Engine behind one key

One API, one balance, sessions routed on live health.