Completed attempts recorded
First flight.
Real work.
A first look at FW-Kimi-K3 through repeatable coding trials and useful improvements to OpenCode.
Repeated trials count separately
Focused regression checks
10 intervals + 12 SSE cases
3,367,336 tokens on record.
All managed local API workers have stopped. Three timed-out OpenCode sessions retain incomplete database status, and cancelled requests may have unreported usage. This is the final recorded tally, not complete provider accounting or an invoice. Counter snapshot: 2026-10-05T21:58:00Z.
How it holds up under pressure.
0 load run(s) active
API response metrics
Transport completion, not answer quality
Errors include 4 timeouts · 0 rate limits
Latest saved client snapshot
Separate from OpenCode above
| Load run | Status | First content p50 / p95 | Latency p50 / p95 | Evidence |
|---|---|---|---|---|
| load-20261005T213700Z-424987stress prompts · 8 workers · 8,192 max generation/request | deadline stopped0 visible · 24 reasoning-only8 client stops · 0 timeouts | 2.0s / 6.9s | 4m 43.0s / 5m 27.8s | Metrics (archived privately) |
| load-20261005T214046Z-633bbdstress prompts · 8 workers · 8,192 max generation/request | deadline stopped0 visible · 24 reasoning-only8 client stops · 0 timeouts | 2.1s / 4.6s | 4m 44.6s / 5m 32.7s | Metrics (archived privately) |
| load-20261005T214252Z-8935fbstress prompts · 16 workers · 8,192 max generation/request | deadline stopped1 visible · 31 reasoning-only16 client stops · 0 timeouts | 3.6s / 5.4s | 5m 19.0s / 5m 20.1s | Metrics (archived privately) |
| load-20261005T214627Z-59fdbbcode prompts · 4 workers · 4,096 max generation/request | completed4 visible · 0 reasoning-only0 client stops · 0 timeouts | 0.5s / 0.8s | 22.7s / 55.6s | Metrics (archived privately) |
| load-20261005T214847Z-46fd1astress prompts · 32 workers · 8,192 max generation/request | operator stopped0 visible · 0 reasoning-only28 client stops · 4 timeouts | 5.0s / 25.6s | 1m 57.2s / 1m 57.2s | Metrics (archived privately) |
5 responses with visible output · 79 reasoning-only streams
Among completed streams. Visible output is not automatically a correct or complete answer.
292,591 prompt tokens · 658,182 reported generation tokens (may include reasoning). Provider-reported cache detail (0) and reasoning detail (0) are not added again. A zero reasoning detail count does not establish that reasoning used no generation tokens; observed reasoning text is counted separately as stream behavior, with no token split inferred from character counts. Usage reported for 84 requests; missing for 64 finished requests. Requests still in flight may not have usage yet.
Raw API cost field: 11.825553300000001 across 84 requests. Unit and currency are unspecified by the response; this is not a verified billed charge or a US dollar total.
Client stops: 28 operator cancellations · 32 scheduled cutoff cancellations · 0 awaiting classification. Stops are excluded from errors. Older probe records may report both kinds as “Shared deadline reached”; preserved intervention evidence supplies the final classification. Other timeouts are client-observed and do not by themselves establish provider failure.
First content measures time to the first text or reasoning delta. Percentiles come from each probe’s recorded timing samples; active requests have not finished contributing. A generation token limit can stop a stream before a final answer appears. This load test measures transport and throughput, not coding correctness.
Evidence, task by task.
Independent verification.
Original tests protected.
| Task | Result | Tests passed | Elapsed | Input / reported output* | Evidence |
|---|---|---|---|---|---|
| intervalsbenchmark · Trial 1 | passed | 10/10Protected tests unchanged | 1m 00.0s | 10,630 / 2,841+ 13,824 cached input | Test log (archived privately) |
| OpenCode debug-config array URL redactionsource | passed | 10/10Protected tests unchanged | 1m 57.6s | 19,100 / 5,836+ 72,192 cached input | Test log (archived privately) |
| ssebenchmark · Trial 1 | passed | 12/12Protected tests unchanged | 2m 31.6s | 25,567 / 7,881+ 102,912 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 2 | passed | 10/10Protected tests unchanged | 51.0s | 8,942 / 2,581+ 13,824 cached input | Test log (archived privately) |
| ssebenchmark · Trial 2 | passed | 12/12Protected tests unchanged | 3m 09.7s | 18,715 / 8,432+ 93,696 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 3 | passed | 10/10Protected tests unchanged | 1m 25.0s | 9,143 / 2,898+ 15,360 cached input | Test log (archived privately) |
| ssebenchmark · Trial 3 | passed | 12/12Protected tests unchanged | 3m 42.6s | 12,110 / 6,904+ 23,040 cached input | Test log (archived privately) |
| OpenCode narrow-terminal middle truncationsource | passed | 10/10Protected tests unchanged | 1m 19.6s | 15,340 / 3,931+ 52,224 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 3 | passed | 10/10Protected tests unchanged | 1m 09.9s | 8,586 / 2,527+ 13,824 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 1 | passed | 10/10Protected tests unchanged | 1m 19.8s | 10,599 / 2,871+ 13,824 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 2 | passed | 10/10Protected tests unchanged | 1m 33.3s | 10,295 / 3,407+ 15,360 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 4 | passed | 10/10Protected tests unchanged | 1m 33.8s | 10,250 / 3,403+ 15,360 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 5 | passed | 10/10Protected tests unchanged | 44.3s | 6,736 / 1,516+ 13,824 cached input | Test log (archived privately) |
| ssebenchmark · Trial 2 | passed | 12/12Protected tests unchanged | 2m 15.6s | 13,338 / 4,976+ 16,896 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 6 | passed | 10/10Protected tests unchanged | 1m 10.6s | 8,578 / 2,439+ 13,824 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 7 | passed | 10/10Protected tests unchanged | 1m 21.5s | 9,167 / 2,684+ 13,824 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 8 | passed | 10/10Protected tests unchanged | 1m 39.6s | 8,402 / 3,140+ 15,360 cached input | Test log (archived privately) |
| ssebenchmark · Trial 3 | passed | 12/12Protected tests unchanged | 4m 25.0s | 26,867 / 9,073+ 109,056 cached input | Test log (archived privately) |
| ssebenchmark · Trial 1 | passed | 12/12Protected tests unchanged | 4m 54.0s | 26,583 / 10,077+ 135,168 cached input | Test log (archived privately) |
| ssebenchmark · Trial 7 | passed | 12/12Protected tests unchanged | 2m 55.5s | 13,069 / 5,586+ 18,432 cached input | Test log (archived privately) |
| ssebenchmark · Trial 4 | passed | 12/12Protected tests unchanged | 5m 22.7s | 27,006 / 10,955+ 118,272 cached input | Test log (archived privately) |
| ssebenchmark · Trial 6 | passed | 12/12Protected tests unchanged | 4m 35.9s | 25,322 / 8,923+ 125,952 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 10 | passed | 10/10Protected tests unchanged | 1m 17.6s | 9,980 / 2,385+ 13,824 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 9 | passed | 10/10Protected tests unchanged | 1m 49.3s | 10,320 / 3,411+ 15,360 cached input | Test log (archived privately) |
| ssebenchmark · Trial 5 | passed | 12/12Protected tests unchanged | 4m 58.8s | 22,193 / 9,820+ 92,160 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 11 | passed | 10/10Protected tests unchanged | 1m 28.1s | 10,327 / 2,613+ 13,824 cached input | Test log (archived privately) |
| ssebenchmark · Trial 8 | passed | 12/12Protected tests unchanged | 3m 54.1s | 22,980 / 7,199+ 89,088 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 12 | passed | 10/10Protected tests unchanged | 1m 13.1s | 9,048 / 2,040+ 13,824 cached input | Test log (archived privately) |
| ssebenchmark · Trial 9 | passed | 12/12Protected tests unchanged | 3m 08.7s | 11,519 / 5,753+ 21,504 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 13 | passed | 10/10Protected tests unchanged | 1m 21.9s | 8,452 / 2,372+ 13,824 cached input | Test log (archived privately) |
| ssebenchmark · Trial 11 | passed | 12/12Protected tests unchanged | 2m 08.1s | 12,615 / 3,835+ 15,360 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 14 | passed | 10/10Protected tests unchanged | 1m 48.5s | 9,931 / 3,203+ 15,360 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 15 | passed | 10/10Protected tests unchanged | 1m 42.0s | 9,550 / 2,872+ 13,824 cached input | Test log (archived privately) |
| ssebenchmark · Trial 10 | passed | 12/12Protected tests unchanged | 4m 56.1s | 21,432 / 8,633+ 86,016 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 16 | passed | 10/10Protected tests unchanged | 1m 51.0s | 9,355 / 2,904+ 15,360 cached input | Test log (archived privately) |
| ssebenchmark · Trial 12 | passed | 12/12Protected tests unchanged | 4m 03.1s | 15,948 / 6,822+ 84,480 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 17 | passed | 10/10Protected tests unchanged | 1m 46.7s | 9,344 / 2,740+ 13,824 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 18 | passed | 10/10Protected tests unchanged | 1m 37.8s | 9,964 / 2,418+ 13,824 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 19 | passed | 10/10Protected tests unchanged | 1m 38.5s | 9,931 / 2,393+ 13,824 cached input | Test log (archived privately) |
| ssebenchmark · Trial 14 | passed | 12/12Protected tests unchanged | 5m 22.2s | 18,647 / 8,582+ 95,232 cached input | Test log (archived privately) |
| intervalsbenchmark · Trial 20 | passed | 10/10Protected tests unchanged | 2m 00.2s | 9,235 / 2,870+ 15,360 cached input | Test log (archived privately) |
| ssebenchmark · Trial 16 | passed | 12/12Protected tests unchanged | 5m 46.9s | 19,519 / 8,710+ 82,944 cached input | Test log (archived privately) |
| ssebenchmark · Trial 13 | timed out | UnverifiedProtected tests unchanged | 8m 00.4s | 29,138 / 12,453+ 159,744 cached input | Test log (archived privately) |
| ssebenchmark · Trial 15 | timed out | 12/12Protected tests unchanged | 8m 00.5s | 18,415 / 10,634+ 87,552 cached input | Test log (archived privately) |
| ssebenchmark · Trial 17 | failed | UnverifiedProtected tests unchanged | 5m 52.9s | 5,520 / 8,694+ 7,680 cached input | Test log (archived privately) |
| ssebenchmark · Trial 18 | passed | 12/12Protected tests unchanged | 7m 39.3s | 22,112 / 9,972+ 76,800 cached input | Test log (archived privately) |
| ssebenchmark · Trial 20 | passed | 12/12Protected tests unchanged | 6m 11.0s | 24,073 / 9,275+ 107,520 cached input | Test log (archived privately) |
| ssebenchmark · Trial 19 | timed out | 12/12Protected tests unchanged | 8m 00.4s | 21,732 / 12,008+ 133,632 cached input | Test log (archived privately) |
| Reusable bounded async task poolsource | failed | UnverifiedProtected tests unchanged | 5m 37.8s | 4,261 / 8,428+ 7,680 cached input | Test log (archived privately) |
| Reusable bounded async task pool: focused retrysource | passed | 16/16Protected tests unchanged | 1m 38.9s | 7,617 / 2,721+ 19,968 cached input | Test log (archived privately) |
* Per-task tokens prefer matching local session counters when available; otherwise they use partial emitted step events. Elapsed times include local execution and verification; parallel task times should not be added to infer campaign duration.
Code worth keeping.
Python Utilities
passed20 model assertions passed · 16 independent tests passed
- Verification report (archived privately)
- Source: utilities.py (archived privately)
- Test log (archived privately)
- Independent review log (archived privately)
Standalone deliverable; separate from OpenCode fixes and the 22-case benchmark.
Task Pool
passed16 protected tests passed · 7 independent tests passed
Guided retry passed; the original attempt failed.
- Verification report (archived privately)
- Source: task-pool.js (archived privately)
- Independent review log (archived privately)
Standalone deliverable; separate from OpenCode fixes and the 22-case benchmark.
Thirty minutes. Measured.
- 3:28:14 PMProvider campaign begins · 21:28:14 UTC
- 3:56:28 PMLocal harness cutoff · 21:56:28 UTC
- 3:58:14 PMProvider window ends · 21:58:14 UTC
Local cutoff is a client stop; server cancellation is controlled by the provider.
Keep the receipts.
- opencode-debug-redaction.patch (archived privately)
- opencode-narrow-terminal.patch (archived privately)
Run summaries (5)
- run-20261005T212828Z-418892 (archived privately) completed
- run-20261005T213221Z-762896 (archived privately) completed
- run-20261005T213427Z-a8d84c (archived privately) completed
- run-20261005T214409Z-c19026 (archived privately) completed
- run-20261005T215317Z-01be86 (archived privately) completed
The scorecard is public. Raw logs, run summaries and package source remain in the private archive. Explore K3 Workbench ↗
What this proves—and its limits.
The benchmark has 22 distinct acceptance cases; repeated attempts demonstrate consistency, not additional unique cases. Passing test executions include recorded source regression tests and standalone library acceptance tests; those are separate from the 22-case benchmark. Counts come from completed, verified attempts in the saved summaries.
Configured context/output limits are 32,768 / 8,192 tokens: testing assumptions, not independently confirmed model specifications. Full OpenCode build, lint, and typecheck were not run for these source fixes. Reported output counters may include reasoning that the provider did not identify separately. Character counts are never converted into a token split. The direct API raw cost field is shown separately with unspecified unit and currency; it is not a verified billed charge. Zero-valued OpenCode cost fields are not treated as free usage.