AGP Picks View all

Floatboat says its harness beats Claude Opus 4.8 on five benchmarks

Aug. 28, 2026
By AI, Created 07:40 UTC, Aug 28, 2026, AGP -

Floatboat says the same DeepSeek-V4-Flash model, when run on Floatboat Harness, outperformed Claude Opus 4.8 across five third-party benchmarks while costing 57.1 times less. The claim puts the spotlight on agent infrastructure, not just model quality, as a driver of real-world performance.

Why it matters: - Floatboat is arguing that agent performance depends as much on the execution system as on the model itself. - The company says its harness turned a low-cost model into a benchmark leader while keeping model weights and unit costs unchanged. - The claim matters because long-horizon work is where agents fail most often, and that is where runtime, tools, persistence, and loop design have the most impact.

What happened: - Floatboat says the same DeepSeek-V4-Flash model outperformed Claude Opus 4.8 across five benchmarks when paired with Floatboat Harness. - The company released the results on August 7 with complete evaluation data across five third-party benchmarks. - The source announcement was dated August 28, 2026, and said the benchmark set showed the cheapest model could beat top-tier overseas models when run in Floatboat. - The full report is available here.

The details: - DeepSeek-V4-Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens. - Under a typical 3:1 input-output ratio, Floatboat calculates a blended cost of $0.175 per million tokens. - Claude Opus 4.8 sits at a blended $10 per million tokens, or 57.1 times more expensive. - On DeepSWE, DeepSeek’s official Harness scored 54.4, below Opus 4.8 at 58.0. - On Floatboat, the same model reached 67.25 on DeepSWE. - On Terminal Bench 2.1, DeepSeek’s official Harness tied Opus 4.8 at 82.7 to 82.7. - Floatboat says the same DeepSeek-V4-Flash model won the other four benchmarks as well. - Both sides used the DeepSeek-V4-Flash 0731 model base. - DeepSeek’s public benchmarks used the minimalist mode of DeepSeek Harness, configured at top_p=0.95 and temperature=1.0. - Floatboat used the Floatboat Evaluation Harness, with model inference for most benchmarks provided by DeepSeek’s official API. - All tasks ran independently in isolated sandbox environments with identical input parameters. - Floatboat says the cheapest model was chosen deliberately so the result could be attributed only to the execution system. - The same-base Harness deltas rose in order from 1.9% to 9.6% to 12.6% to 19.9% to 23.6%. - Floatboat says the gains increased monotonically as task horizon lengthened. - On BrowseComp, Floatboat’s 87.80 beat GPT-5.6 Terra at 87.5, Claude Sonnet 5 at 84.7, Opus 4.8 at 84.3, and GPT-5.6 Luna at 83.3. - AOE Tech Labs also launched a DeepSeek-based edition priced at ¥5 / $1 for the first month, built on the same DeepSeek-V4-Flash base model.

Between the lines: - The central claim is not that DeepSeek-V4-Flash became a better model, but that the surrounding system unlocked more of its capability. - Floatboat is trying to redefine the benchmark conversation around harness quality, not just model rankings. - The company’s HLR metric, or Harness Leverage Ratio, is meant to show how much improvement comes from changing the harness versus changing the model. - On DeepSWE, Floatboat says HLR equals 12.85 divided by 3.6, or about 3.57. - The broader argument is that long-horizon tasks reward systems that can look back, persist state, and recover from failure instead of just generating the next step. - Floatboat says the user-facing desktop environment is stronger than the evaluation environment, which implies benchmark scores may still understate the product’s full runtime advantage.

What’s next: - Floatboat says every number in the announcement can be recomputed from the cited setup. - The company points readers to floatboat.ai for the entry point and the benchmark report. - Floatboat says the same benchmark set run in the full client environment produces higher scores than the figures disclosed in the announcement. - The company is positioning its runtime, context persistence, and multi-step automation stack as the next battleground for agent performance.

The bottom line: - Floatboat’s message is simple: in agent systems, the harness can matter as much as the model, and sometimes more.

Disclaimer: This article was produced by AGP Wire with the assistance of artificial intelligence based on original source content and has been refined to improve clarity, structure, and readability. This content is provided on an “as is” basis. While care has been taken in its preparation, it may contain inaccuracies or omissions, and readers should consult the original source and independently verify key information where appropriate. This content is for informational purposes only and does not constitute legal, financial, investment, or other professional advice.

Sign up for:

New Products Launch Guide

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

New Products Launch Guide

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.