K3 / FIELD NOTESALPHA TEST · OCTOBER 05, 2026
Saved campaign results

First flight.
Real work.

A first look at FW-Kimi-K3 through repeatable coding trials and useful improvements to OpenCode.

FW-Kimi-K3 × OpenCode 2.0.23 × macOS

Tasks passed45 / 50

Completed attempts recorded

Passing test executions518

Repeated trials count separately

OpenCode fixes verified2

Focused regression checks

Unique benchmark cases22

10 intervals + 12 SSE cases

OPENCODE TOKEN TALLY · FINAL RECORDED

3,367,336 tokens on record.

All managed local API workers have stopped. Three timed-out OpenCode sessions retain incomplete database status, and cancelled requests may have unreported usage. This is the final recorded tally, not complete provider accounting or an invoice. Counter snapshot: 2026-10-05T21:58:00Z.

722,989Uncached input
278,907Reported output
2,365,440Cached input read
0Cache writes
Local session counters · Reasoning reported: 0 · Verified billed charge: unknown Combined recorded usage: 4,318,109 tokens across OpenCode + direct API.
DIRECT API LOAD · SEPARATE ACCOUNTING

How it holds up under pressure.

0 load run(s) active
API response metrics

Completed streams84

Transport completion, not answer quality

Errors / client stops4 / 60

Errors include 4 timeouts · 0 rate limits

Requests in flight0

Latest saved client snapshot

Direct API reported tokens950,773

Separate from OpenCode above

Direct API response timing by load run
Load runStatusFirst content p50 / p95Latency p50 / p95Evidence
load-20261005T213700Z-424987stress prompts · 8 workers · 8,192 max generation/request deadline stopped0 visible · 24 reasoning-only8 client stops · 0 timeouts2.0s / 6.9s 4m 43.0s / 5m 27.8sMetrics (archived privately)
load-20261005T214046Z-633bbdstress prompts · 8 workers · 8,192 max generation/request deadline stopped0 visible · 24 reasoning-only8 client stops · 0 timeouts2.1s / 4.6s 4m 44.6s / 5m 32.7sMetrics (archived privately)
load-20261005T214252Z-8935fbstress prompts · 16 workers · 8,192 max generation/request deadline stopped1 visible · 31 reasoning-only16 client stops · 0 timeouts3.6s / 5.4s 5m 19.0s / 5m 20.1sMetrics (archived privately)
load-20261005T214627Z-59fdbbcode prompts · 4 workers · 4,096 max generation/request completed4 visible · 0 reasoning-only0 client stops · 0 timeouts0.5s / 0.8s 22.7s / 55.6sMetrics (archived privately)
load-20261005T214847Z-46fd1astress prompts · 32 workers · 8,192 max generation/request operator stopped0 visible · 0 reasoning-only28 client stops · 4 timeouts5.0s / 25.6s 1m 57.2s / 1m 57.2sMetrics (archived privately)

5 responses with visible output · 79 reasoning-only streams
Among completed streams. Visible output is not automatically a correct or complete answer.

292,591 prompt tokens · 658,182 reported generation tokens (may include reasoning). Provider-reported cache detail (0) and reasoning detail (0) are not added again. A zero reasoning detail count does not establish that reasoning used no generation tokens; observed reasoning text is counted separately as stream behavior, with no token split inferred from character counts. Usage reported for 84 requests; missing for 64 finished requests. Requests still in flight may not have usage yet.

Raw API cost field: 11.825553300000001 across 84 requests. Unit and currency are unspecified by the response; this is not a verified billed charge or a US dollar total.

Client stops: 28 operator cancellations · 32 scheduled cutoff cancellations · 0 awaiting classification. Stops are excluded from errors. Older probe records may report both kinds as “Shared deadline reached”; preserved intervention evidence supplies the final classification. Other timeouts are client-observed and do not by themselves establish provider failure.

First content measures time to the first text or reasoning delta. Percentiles come from each probe’s recorded timing samples; active requests have not finished contributing. A generation token limit can stop a stream before a final answer appears. This load test measures transport and throughput, not coding correctness.

THE FLIGHT LOG

Evidence, task by task.

Independent verification.
Original tests protected.

Completed task results and partial reported token use
TaskResultTests passedElapsedInput / reported output*Evidence
intervalsbenchmark · Trial 1 passed10/10Protected tests unchanged 1m 00.0s10,630 / 2,841+ 13,824 cached input Test log (archived privately)
OpenCode debug-config array URL redactionsource passed10/10Protected tests unchanged 1m 57.6s19,100 / 5,836+ 72,192 cached input Test log (archived privately)
ssebenchmark · Trial 1 passed12/12Protected tests unchanged 2m 31.6s25,567 / 7,881+ 102,912 cached input Test log (archived privately)
intervalsbenchmark · Trial 2 passed10/10Protected tests unchanged 51.0s8,942 / 2,581+ 13,824 cached input Test log (archived privately)
ssebenchmark · Trial 2 passed12/12Protected tests unchanged 3m 09.7s18,715 / 8,432+ 93,696 cached input Test log (archived privately)
intervalsbenchmark · Trial 3 passed10/10Protected tests unchanged 1m 25.0s9,143 / 2,898+ 15,360 cached input Test log (archived privately)
ssebenchmark · Trial 3 passed12/12Protected tests unchanged 3m 42.6s12,110 / 6,904+ 23,040 cached input Test log (archived privately)
OpenCode narrow-terminal middle truncationsource passed10/10Protected tests unchanged 1m 19.6s15,340 / 3,931+ 52,224 cached input Test log (archived privately)
intervalsbenchmark · Trial 3 passed10/10Protected tests unchanged 1m 09.9s8,586 / 2,527+ 13,824 cached input Test log (archived privately)
intervalsbenchmark · Trial 1 passed10/10Protected tests unchanged 1m 19.8s10,599 / 2,871+ 13,824 cached input Test log (archived privately)
intervalsbenchmark · Trial 2 passed10/10Protected tests unchanged 1m 33.3s10,295 / 3,407+ 15,360 cached input Test log (archived privately)
intervalsbenchmark · Trial 4 passed10/10Protected tests unchanged 1m 33.8s10,250 / 3,403+ 15,360 cached input Test log (archived privately)
intervalsbenchmark · Trial 5 passed10/10Protected tests unchanged 44.3s6,736 / 1,516+ 13,824 cached input Test log (archived privately)
ssebenchmark · Trial 2 passed12/12Protected tests unchanged 2m 15.6s13,338 / 4,976+ 16,896 cached input Test log (archived privately)
intervalsbenchmark · Trial 6 passed10/10Protected tests unchanged 1m 10.6s8,578 / 2,439+ 13,824 cached input Test log (archived privately)
intervalsbenchmark · Trial 7 passed10/10Protected tests unchanged 1m 21.5s9,167 / 2,684+ 13,824 cached input Test log (archived privately)
intervalsbenchmark · Trial 8 passed10/10Protected tests unchanged 1m 39.6s8,402 / 3,140+ 15,360 cached input Test log (archived privately)
ssebenchmark · Trial 3 passed12/12Protected tests unchanged 4m 25.0s26,867 / 9,073+ 109,056 cached input Test log (archived privately)
ssebenchmark · Trial 1 passed12/12Protected tests unchanged 4m 54.0s26,583 / 10,077+ 135,168 cached input Test log (archived privately)
ssebenchmark · Trial 7 passed12/12Protected tests unchanged 2m 55.5s13,069 / 5,586+ 18,432 cached input Test log (archived privately)
ssebenchmark · Trial 4 passed12/12Protected tests unchanged 5m 22.7s27,006 / 10,955+ 118,272 cached input Test log (archived privately)
ssebenchmark · Trial 6 passed12/12Protected tests unchanged 4m 35.9s25,322 / 8,923+ 125,952 cached input Test log (archived privately)
intervalsbenchmark · Trial 10 passed10/10Protected tests unchanged 1m 17.6s9,980 / 2,385+ 13,824 cached input Test log (archived privately)
intervalsbenchmark · Trial 9 passed10/10Protected tests unchanged 1m 49.3s10,320 / 3,411+ 15,360 cached input Test log (archived privately)
ssebenchmark · Trial 5 passed12/12Protected tests unchanged 4m 58.8s22,193 / 9,820+ 92,160 cached input Test log (archived privately)
intervalsbenchmark · Trial 11 passed10/10Protected tests unchanged 1m 28.1s10,327 / 2,613+ 13,824 cached input Test log (archived privately)
ssebenchmark · Trial 8 passed12/12Protected tests unchanged 3m 54.1s22,980 / 7,199+ 89,088 cached input Test log (archived privately)
intervalsbenchmark · Trial 12 passed10/10Protected tests unchanged 1m 13.1s9,048 / 2,040+ 13,824 cached input Test log (archived privately)
ssebenchmark · Trial 9 passed12/12Protected tests unchanged 3m 08.7s11,519 / 5,753+ 21,504 cached input Test log (archived privately)
intervalsbenchmark · Trial 13 passed10/10Protected tests unchanged 1m 21.9s8,452 / 2,372+ 13,824 cached input Test log (archived privately)
ssebenchmark · Trial 11 passed12/12Protected tests unchanged 2m 08.1s12,615 / 3,835+ 15,360 cached input Test log (archived privately)
intervalsbenchmark · Trial 14 passed10/10Protected tests unchanged 1m 48.5s9,931 / 3,203+ 15,360 cached input Test log (archived privately)
intervalsbenchmark · Trial 15 passed10/10Protected tests unchanged 1m 42.0s9,550 / 2,872+ 13,824 cached input Test log (archived privately)
ssebenchmark · Trial 10 passed12/12Protected tests unchanged 4m 56.1s21,432 / 8,633+ 86,016 cached input Test log (archived privately)
intervalsbenchmark · Trial 16 passed10/10Protected tests unchanged 1m 51.0s9,355 / 2,904+ 15,360 cached input Test log (archived privately)
ssebenchmark · Trial 12 passed12/12Protected tests unchanged 4m 03.1s15,948 / 6,822+ 84,480 cached input Test log (archived privately)
intervalsbenchmark · Trial 17 passed10/10Protected tests unchanged 1m 46.7s9,344 / 2,740+ 13,824 cached input Test log (archived privately)
intervalsbenchmark · Trial 18 passed10/10Protected tests unchanged 1m 37.8s9,964 / 2,418+ 13,824 cached input Test log (archived privately)
intervalsbenchmark · Trial 19 passed10/10Protected tests unchanged 1m 38.5s9,931 / 2,393+ 13,824 cached input Test log (archived privately)
ssebenchmark · Trial 14 passed12/12Protected tests unchanged 5m 22.2s18,647 / 8,582+ 95,232 cached input Test log (archived privately)
intervalsbenchmark · Trial 20 passed10/10Protected tests unchanged 2m 00.2s9,235 / 2,870+ 15,360 cached input Test log (archived privately)
ssebenchmark · Trial 16 passed12/12Protected tests unchanged 5m 46.9s19,519 / 8,710+ 82,944 cached input Test log (archived privately)
ssebenchmark · Trial 13 timed outUnverifiedProtected tests unchanged 8m 00.4s29,138 / 12,453+ 159,744 cached input Test log (archived privately)
ssebenchmark · Trial 15 timed out12/12Protected tests unchanged 8m 00.5s18,415 / 10,634+ 87,552 cached input Test log (archived privately)
ssebenchmark · Trial 17 failedUnverifiedProtected tests unchanged 5m 52.9s5,520 / 8,694+ 7,680 cached input Test log (archived privately)
ssebenchmark · Trial 18 passed12/12Protected tests unchanged 7m 39.3s22,112 / 9,972+ 76,800 cached input Test log (archived privately)
ssebenchmark · Trial 20 passed12/12Protected tests unchanged 6m 11.0s24,073 / 9,275+ 107,520 cached input Test log (archived privately)
ssebenchmark · Trial 19 timed out12/12Protected tests unchanged 8m 00.4s21,732 / 12,008+ 133,632 cached input Test log (archived privately)
Reusable bounded async task poolsource failedUnverifiedProtected tests unchanged 5m 37.8s4,261 / 8,428+ 7,680 cached input Test log (archived privately)
Reusable bounded async task pool: focused retrysource passed16/16Protected tests unchanged 1m 38.9s7,617 / 2,721+ 19,968 cached input Test log (archived privately)

* Per-task tokens prefer matching local session counters when available; otherwise they use partial emitted step events. Elapsed times include local execution and verification; parallel task times should not be added to infer campaign duration.

THE DELIVERABLES

Code worth keeping.

GENERATED PROJECT

Python Utilities

passed

20 model assertions passed · 16 independent tests passed

  • Verification report (archived privately)
  • Source: utilities.py (archived privately)
  • Test log (archived privately)
  • Independent review log (archived privately)

Standalone deliverable; separate from OpenCode fixes and the 22-case benchmark.

GENERATED PROJECT

Task Pool

passed

16 protected tests passed · 7 independent tests passed

Guided retry passed; the original attempt failed.

  • Verification report (archived privately)
  • Source: task-pool.js (archived privately)
  • Independent review log (archived privately)

Standalone deliverable; separate from OpenCode fixes and the 22-case benchmark.

THE WINDOW · DENVER / MDT

Thirty minutes. Measured.

  1. 3:28:14 PMProvider campaign begins · 21:28:14 UTC
  2. 3:56:28 PMLocal harness cutoff · 21:56:28 UTC
  3. 3:58:14 PMProvider window ends · 21:58:14 UTC

Local cutoff is a client stop; server cancellation is controlled by the provider.

REVIEWABLE ARTIFACTS

Keep the receipts.

  • opencode-debug-redaction.patch (archived privately)
  • opencode-narrow-terminal.patch (archived privately)
Run summaries (5)
  • run-20261005T212828Z-418892 (archived privately) completed
  • run-20261005T213221Z-762896 (archived privately) completed
  • run-20261005T213427Z-a8d84c (archived privately) completed
  • run-20261005T214409Z-c19026 (archived privately) completed
  • run-20261005T215317Z-01be86 (archived privately) completed

The scorecard is public. Raw logs, run summaries and package source remain in the private archive. Explore K3 Workbench ↗

What this proves—and its limits.

The benchmark has 22 distinct acceptance cases; repeated attempts demonstrate consistency, not additional unique cases. Passing test executions include recorded source regression tests and standalone library acceptance tests; those are separate from the 22-case benchmark. Counts come from completed, verified attempts in the saved summaries.

Configured context/output limits are 32,768 / 8,192 tokens: testing assumptions, not independently confirmed model specifications. Full OpenCode build, lint, and typecheck were not run for these source fixes. Reported output counters may include reasoning that the provider did not identify separately. Character counts are never converted into a token split. The direct API raw cost field is shown separately with unspecified unit and currency; it is not a verified billed charge. Zero-valued OpenCode cost fields are not treated as free usage.