Tech9 min read

AWS Strands Harness vs Claude Code & Codex: Same-Model Benchmarks

IkesanContents

The AWS Strands Agents team released “Strands harness” on September 21, 2026.
It provides a pre-assembled harness for AI agents out of the box, with support for both Python and TypeScript under the Apache 2.0 license.
The underlying LLM can be swapped with a single string passed to the model parameter.
In Japan, Publickey and ITmedia NEWS covered the launch on September 29, 2026.

The headline claim is “28% cheaper using the same models as Claude Code and Codex.”
The charts in the official blog post are embedded HTML elements with raw underlying numbers exposed.
Extracting those numbers lets us see how Strands harness fares against Claude Code and Codex when running on identical models.

A Harness Is Everything Outside the LLM

An AI agent does not work with an LLM alone.
When the model outputs a command to run, the system must execute it and return the result. When context grows too long, older parts need summarizing to free up space. When resuming, past sessions need loading.
Everything handling these tasks outside the LLM itself is referred to as the harness.
Both Claude Code and the Codex CLI are built by pairing their respective harnesses with a model.

The official blog begins by addressing a common issue: an agent setup that worked well inside Claude Code or Codex fails to behave the same way when reconstructed from scratch.
Strands harness distributes this “pre-assembled harness” as a library, engineered to run in cloud environments as well.
The team notes it was designed not exclusively for coding, but as a general-purpose agent harness.

The foundation is the “Strands Harness SDK” from the same team, hosted on GitHub at strands-agents/harness-sdk under harness-py and harness-ts.
While the SDK offers low-level building blocks, Strands harness is the ready-to-run package assembled with sensible defaults.

Out-of-the-Box Features

Calling create_harness() with no arguments returns an agent with the following capabilities enabled by default (based on the configuration reference):

FeatureDefault Behavior
ModelClaude Opus 5 on Amazon Bedrock (bedrock/global.anthropic.claude-opus-5)
Reasoning Efforteffort="auto", using the provider’s recommended level
Built-in Toolsshell, read, write, edit, web_fetch, web_search, programmatic_tool_caller, subagent
Context ManagementTruncates tool outputs over ~1,500 tokens; compresses via summarization above 85% context; recovers in-loop if exceeded
Prompt CachingEnabled where supported by the provider
SessionsResumes past conversations by passing a session ID. Stored in ./.agent/sessions
Long-term MemoryFile-based memory enabled. Stored in ./.agent/memory
SkillsLoads Agent Skills from ./.agent/skills
SubagentsDelegates subtasks to the built-in generalist subagent
ToDoTracks multi-step tasks via checklists

The programmatic_tool_caller allows the agent to call tools from within model-generated code.

The blog post explains that most of the token efficiency and benchmark accuracy advantages stem from these context management defaults.
Excessively long tool outputs are truncated and saved to disk, while conversation history is compressed before context limits are reached.

Swapping LLMs with a Single String

Models are specified as a provider/model-name string.
According to the model configuration guide, seven provider prefixes are supported: bedrock, bedrock-mantle, anthropic, openai, google, ollama, and litellm.
Strings without a prefix are treated as Amazon Bedrock model IDs.

# pip install strands-harness
from strands_harness import create_harness

agent = create_harness(model="openai/gpt-5.6-sol")
agent("Research the top three vector databases, compare pricing and limits, and write it up in comparison.md")

In TypeScript, this is written as createHarness({ model: "openai/gpt-5.6-sol" }).
Replacing the string with anthropic/claude-opus-5 or google/gemini-3.5-flash runs the same agent on another model, while local Ollama instances can be targeted using ollama/llama3.1.
Routing through LiteLLM enables connections to providers without native aliases.

Reasoning intensity is adjusted via effort with "low", "medium", "high", or "off".
Differences in reasoning parameters across providers are normalized by the library.
Specifying a level unsupported by a provider raises an error at harness creation time rather than mid-request.

Feature availability varies depending on the chosen model.
Native web_search is limited to providers including OpenAI, Anthropic, and Google. Prompt caching is explicitly managed by Strands harness only for Bedrock and direct Anthropic connections; OpenAI, Gemini, Bedrock Mantle, and LiteLLM handle caching automatically on the provider side, while Ollama does not support prompt caching.

Benchmark Numbers from Official Data

The official blog post presents two charts: one averaging six benchmark suites, and another focusing specifically on Terminal-Bench 2.1 for terminal-based tasks.
Both tables below are transcribed directly from raw numbers inside the embedded chart HTML.
AWS reports that measurements were run distributed across Amazon EC2 using the Harbor evaluation framework.

Average Across Six Benchmarks

The benchmark suite includes ALFWorld, ContextBench, GAIA, WebShop, τ³-bench, and Terminal-Bench 2.1.
Scores reflect average accuracy across all six, while costs denote average dollar spend per task.
The five competitors tested were Claude Code, Codex, oh-my-pi, OpenCode, and DeepSeek Harness.
The table below pairs Strands harness against Claude Code and Codex when configured with identical models:

ModelCompetitorCompetitor ScoreCompetitor CostStrands ScoreStrands Cost
Claude Fable 5Claude Code80.94%$1.87281.11%$0.829
Claude Opus 5Claude Code84.60%$1.06384.55%$0.634
Claude Opus 4.8Claude Code77.32%$0.98377.14%$0.421
Claude Sonnet 5Claude Code71.55%$0.97475.61%$0.311
GPT-5.6 SolCodex76.59%$0.62876.72%$0.293
GPT-5.6 TerraCodex59.06%$0.17864.75%$0.143
GPT-5.6 LunaCodex62.66%$0.02861.94%$0.009

For five of the seven pairs, score differences stayed under one percentage point.
The only notable gaps appeared with Claude Sonnet 5 (~4 points) and GPT-5.6 Terra (~6 points), where Strands harness scored higher in both cases.
Across all seven pairings, Strands harness cost less, with expenses reduced to between roughly 60% and 33% of Claude Code.

The headline “28%” refers to the combined average across all five competitors.
Chart footnotes indicate that DeepSeek Harness was ~14% cheaper than Strands harness, but lagged behind on all benchmark scores.
Including DeepSeek Harness in the overall comparison pulled the aggregated savings figure down to 28%.
Compared strictly against Claude Code, the cost reduction is substantially wider.

Terminal-Bench 2.1 with Claude Fable 5

The second chart compares all harnesses running Claude Fable 5 across 89 trials on Terminal-Bench 2.1:

HarnessAccuracyCostTokens
DeepSeek Harness59.55%$40.3021.44M
Strands harness69.66%$56.2928.73M
OpenCode66.29%$73.4237.11M
oh-my-pi69.66%$86.8345.06M
Claude Code61.80%$248.0554.99M

The “77% cheaper than Claude Code” claim comes directly from this dataset.
$56.29 is roughly 23% of $248.05, while Strands harness achieved a 7.86-point lead in accuracy.
While oh-my-pi tied Strands in accuracy at 69.66%, its total cost was about 1.5x higher.

Tokens differed by roughly 1.9x, yet cost diverged by approximately 4.4x.
The exact drivers behind this pricing gap are not detailed in the official blog post.

Caveats When Reading the Benchmark Data

These benchmarks were measured and published directly by AWS, with an academic paper promised for a later date.
The Register noted that while Strands harness is marketed as a general-purpose agent, all of its evaluation targets are dedicated coding agents.

The chart source code also disclosed several caveats regarding data collection:

Note in SourcePractical Impact
Only pairs with complete cost data across all 6 benchmarks were plottedExcluded eve, deepagents, and GPT-6 Astra. Astra was missing cost numbers for most benchmarks
Missing cells were backfilled from a separate report (harbor-report-sep17)Applied to Claude Code + Fable 5 on Terminal-Bench, 4 ContextBench points for oh-my-pi, and Codex + GPT-5.6 Terra on ALFWorld
Costs used measured values when available, falling back to estimatesThe chart does not indicate which individual cells rely on estimates

The Claude Code + Fable 5 row in the six-benchmark average borrows its Terminal-Bench datapoint from a separate run.
While the numbers above reflect published figures, these backfills and estimates provide essential context when evaluating the results.

Default Permissions and Unprompted Shell Execution

Coming from Claude Code, the most jarring default is tool execution permissions.
By default, Strands harness configures no intervention hooks (interventions), so every tool call executes immediately without confirmation.
Because shell and file operations execute directly on the local host environment, an agent initialized without arguments runs commands on your system unprompted.

According to the interventions documentation, four approval modes are available:

SettingBehavior
interventions="ask"Prompts for confirmation before every tool call
interventions="smart"Prompts only when the SDK’s risk classifier flags an action as hazardous
Natural language stringUses guidelines like “confirm before file deletion or network requests” as criteria
Path ending in .cedarEnforces permissions defined in a Cedar policy file

The built-in generalist subagent inherits these settings, so delegated subtasks cannot bypass approval gates.
The production deployment guide also emphasizes locking down approvals for shell and file edits alongside programmatic_tool_caller.
While programmatic_tool_caller executes model-authored code inside an isolated Monty environment, tools invoked from that code are not isolated.
For untrusted inputs, AWS recommends executing inside Docker or SSH sandboxes or disabling the tool altogether.

The shell tool spawns a fresh process for each invocation. Working directories and exported environment variables do not persist across calls.
File tools require absolute paths and reject paths containing ...

Deployment Targets

The library can be deployed anywhere Linux containers run. The blog highlights targets including Modal, Cloudflare Containers, Azure Container Apps, Google Cloud Run, Amazon ECS, and Amazon Bedrock AgentCore.
Because create_harness() returns a standard Strands Agent, deployment workflows written for the SDK apply directly.

When deploying to containers or serverless environments, storage configuration is critical.
Sessions and long-term memory default to ./.agent, which vanishes when ephemeral instances shut down.
Point them to persistent storage using session={"dir": ...} and memory={"dir": ...}.

Installation and the Strands CLI

The library can be installed via pip install strands-harness (Python 3.10+) or npm install @strands-agents/harness.
As of September 29, 2026, releases are 0.1.2 on PyPI and 0.1.1 on npm.

For interactive prototyping, the team provides the Strands CLI.
In the official demo, the CLI was used to select a model and tools, prompt the agent to “add Playwright MCP and measure video loading latency on a blog post,” and test execution.
Running /export afterward generates ready-to-use Python or TypeScript code incorporating the Playwright MCP configuration.
The CLI itself is built directly on top of Strands harness.