Kimchi Inference

    Open-source models. One line to switch.

    A serverless API for production-ready open-source LLMs. OpenAI- and Anthropic-compatible, so your existing client works today. No GPUs to provision, no clusters to manage.

    OpenAI-compatibleAnthropic-compatiblePay per token, per model
    first request · any OpenAI-compatible client
    $ curl https://llm.kimchi.dev/openai/v1/chat/completions \  -H "Content-Type: application/json" \  -H "Authorization: Bearer $KIMCHI_API_KEY" \  -d '{    "model": "kimi-k3",    "messages": [{ "role": "user", "content": "Explain a DaemonSet in one sentence." }]  } // Standard OpenAI chat completions response"model": "kimi-k3", "usage": { "prompt_tokens": 12, "completion_tokens": 38 }

    [ The switch ]

    One API change. That's the whole migration.

    Point any OpenAI- or Anthropic-compatible client at Kimchi. Same request format, same response format, same client code. The migration is this diff:

    your client config · the entire migration
    - base_url = "https://api.your-current-provider.com/v1"+ base_url = "https://llm.kimchi.dev/openai/v1" // that's it. same client, same code, same request format.

    [ Deployment modes ]

    Start serverless. Go self-hosted when control matters.

    The same models, two ways to run them. Move when per-token costs exceed compute costs, or when compliance requires your own cluster.

    Serverless

    Kimchi-hosted

    Kimchi runs the models on our own GPU infrastructure. You pay per token, nothing to manage.

    Live in minutes with an API key
    8 models, 8 price points
    Per-user, per-team, per-key spend visibility
    VPC

    Self-hosted

    Run the same models on your own Kubernetes cluster when scale or compliance demands it.

    Same API, same models, your GPUs
    When per-token beats compute cost, stay serverless
    Hosted Model Deployment in the docs
    Autoscaling on your own hardware
    BRING YOUR OWN KEY

    External providers

    Connect external providers through the same endpoint.  on top of the provider's own token pricing.

    Frontier models when you need them
    One bill, one place for spend visibility
    Per-token surcharge, flat across models

    [ Models & pricing ]

    Eight models. Real prices.

    Every model runs on Kimchi GPU infrastructure. No per-seat charge for inference, no minimum spend. Pick per request, switch anytime.

    ModelBest forContextInput / 1MCached input / 1MOutput / 1M
    deepseek-v4-flashFast inference, cost-efficient tasks1M$0.14$0.07$0.28
    deepseek-v4.1-flashFast inference, cost-efficient tasks and image understanding1M$0.30$0.006$1.20
    glm-5.3Complex reasoning1M$1.40$0.14$4.40
    glm-5.3-flashFast, cost-efficient reasoning and image understanding1M$0.15$0.03$0.50
    kimi-k2.7Agentic coding, image analysis262.1K$0.95$0.19$4.00
    kimi-k3Agentic coding, image and video analysis262.1K$3.00$0.30$15.00
    minimax-m3Orchestration, planning, coding, review, image understanding1M$0.30$0.06$1.20
    nemotron-3-ultra-fp4Fast inference, cost-efficient tasks1M$0.60—$3.60
    Avg. price of proprietary models———$12.10
    Per 1M tokens, per model, through the serverless API. Full table: docs.kimchi.dev/docs/model-apis-pricing

    Enterprise controls backed by Cast AI's security program

    Security, compliance, and control aren't afterthoughts. Every feature ships enterprise-ready.

    ISO 27001GDPRAICPA SOC 2

    One API change. Eight models.

    Create an API key, point your client at llm.kimchi.dev. Same code, same format, running on open-source models tomorrow.