← All posts
    EngineeringJuly 28, 2026·5 min read

    A two-day hackathon on Kimchi inference, at 1/13th the frontier cost

    ~120,700 requests. 1,580 active sessions. ~98% success rate. ~$4.4k total spend - about 1/13th of the modeled frontier-model cost.

    A two-day hackathon on Kimchi inference, at 1/13th the frontier cost

    ~120,700 requests. 1,580 active sessions. ~98% success rate. ~$4.4k total spend.

    That's two days of real load on our inference layer. Last week (20-21 July) we ran a hackathon in Vilnius, and everything - the agents, the tooling, the side projects - ran on open-source models served through Kimchi.dev, our own inference platform.

    Nobody was forced to use our models. People reached for them by default, and they held up. Here's what the telemetry says.

    One number for scale: we modeled the same workload on a frontier-model mix (80% Claude Sonnet 5, 20% Claude Fable 5). That comes out to roughly $58k. Same workload, about 1/13th the cost on our own inference.

    Our inference vs. modeled frontier mix - same workload, about 1/13th the cost

    The model mix

    glm-5.2-fp8 was the workhorse: 57% of requests and $3,673 of the ~$4.4k spend.

    MODEL                   SHARE OF REQUESTS   COST
    glm-5.2-fp8             57%                 $3,673
    kimi-k2.7               24%                 $542
    minimax-m3              10%                 $113
    deepseek-v4-flash       8%                  $7
    nemotron-3-ultra-fp4    ~1%                 $46

    The spread makes sense. glm-5.2-fp8 took the heavy agentic work, kimi-k2.7 handled tight-loop tasks, and deepseek-v4-flash ran 8% of all traffic for $7.

    Caching did 80% of the work

    Our models were presented with ~10 billion tokens of context over the two days. We only computed ~1.9B of it.

    The other ~8.1B (about 80%) was served from cache - repeated agentic context reused instead of reprocessed.

    ~80% of ~10B tokens served from cache - roughly 5x compute avoided
    MODEL          FRESH INPUT   FROM CACHE   HANDLED    % CACHED
    glm-5.2-fp8    1.45B         5.66B        ~7.15B     ~80%
    kimi-k2.7      0.10B         2.12B        ~2.23B     ~96%
    minimax-m3     0.30B         0.24B        ~0.55B     ~44%

    The per-model cache rates tell you what kind of work each model did. kimi-k2.7 reused ~96% of its context - tight agentic loops resending near-identical prompts. minimax-m3 cached ~44%, doing proportionally more new work per request.

    Without caching we'd have processed the full ~10B - roughly 5x the compute and cost. At agentic workloads, caching isn't an optimization detail. It's the difference between a $4.4k bill and a much worse one.

    Failures were rare. Latency wasn't.

    Error rates stayed between 0.8% and 2% across every model, over two full days of traffic.

    Most of those errors weren't server faults. glm-5.2-fp8 is our heaviest model at ~20s average response time, and some clients timed out and cancelled long requests. That accounts for most of the error count - a client timeout tuning item, not a stability problem.

    Genuine server faults were negligible.

    46 autonomous agent runs, up to 3.5 hours each

    Our /ferment harness launched 46 autonomous runs during the hackathon: ~92% completion rate (of resolved runs), zero retry loops, zero step failures.

    46 autonomous runs, ~92% completion, zero retry loops, longest 1-3.5h

    These weren't demos. Runs shipped real product work - autoscaler binpacking changes, a Spot-optimization UI, Terraform tooling. The longest runs went 1 to 3.5 hours unattended, predominantly on glm-5.2-fp8.

    Humans steered runs by choice, not because the agent got stuck. Most runs graded A on internal review.

    Usage looked like real work

    Traffic peaked at ~9,500 requests/hour and followed a workday rhythm both days: morning ramp, midday peak, evening taper. Organic human use, not synthetic load.

    Monday ran hotter per hour than Tuesday (~4,300 vs ~2,700 req/hr), and usage concentrated in the tooling we built rather than scattering across surfaces.

    What we're taking away

    Three things.

    Open-source models on our own inference handled a real two-day, multi-agent workload at ~98% success for ~$4.4k - about 1/13th of the modeled frontier-model cost.

    Caching is a first-order cost lever for agentic workloads - 80% of our context volume never touched compute.

    The remaining friction is latency on our heaviest model. That's a tuning problem, not an architecture problem.

    Run one yourself

    We're now sponsoring company hackathons. If you want CAST AI to sponsor your team's next one - inference included - reach out and let's talk.