AWS Strands Harness vs Claude Code & Codex: Same-Model Benchmarks
Contents
The AWS Strands Agents team released “Strands harness” on September 21, 2026.
It provides a pre-assembled harness for AI agents out of the box, with support for both Python and TypeScript under the Apache 2.0 license.
The underlying LLM can be swapped with a single string passed to the model parameter.
In Japan, Publickey and ITmedia NEWS covered the launch on September 29, 2026.
The headline claim is “28% cheaper using the same models as Claude Code and Codex.”
The charts in the official blog post are embedded HTML elements with raw underlying numbers exposed.
Extracting those numbers lets us see how Strands harness fares against Claude Code and Codex when running on identical models.
A Harness Is Everything Outside the LLM
An AI agent does not work with an LLM alone.
When the model outputs a command to run, the system must execute it and return the result. When context grows too long, older parts need summarizing to free up space. When resuming, past sessions need loading.
Everything handling these tasks outside the LLM itself is referred to as the harness.
Both Claude Code and the Codex CLI are built by pairing their respective harnesses with a model.
The official blog begins by addressing a common issue: an agent setup that worked well inside Claude Code or Codex fails to behave the same way when reconstructed from scratch.
Strands harness distributes this “pre-assembled harness” as a library, engineered to run in cloud environments as well.
The team notes it was designed not exclusively for coding, but as a general-purpose agent harness.
The foundation is the “Strands Harness SDK” from the same team, hosted on GitHub at strands-agents/harness-sdk under harness-py and harness-ts.
While the SDK offers low-level building blocks, Strands harness is the ready-to-run package assembled with sensible defaults.
Out-of-the-Box Features
Calling create_harness() with no arguments returns an agent with the following capabilities enabled by default (based on the configuration reference):
| Feature | Default Behavior |
|---|---|
| Model | Claude Opus 5 on Amazon Bedrock (bedrock/global.anthropic.claude-opus-5) |
| Reasoning Effort | effort="auto", using the provider’s recommended level |
| Built-in Tools | shell, read, write, edit, web_fetch, web_search, programmatic_tool_caller, subagent |
| Context Management | Truncates tool outputs over ~1,500 tokens; compresses via summarization above 85% context; recovers in-loop if exceeded |
| Prompt Caching | Enabled where supported by the provider |
| Sessions | Resumes past conversations by passing a session ID. Stored in ./.agent/sessions |
| Long-term Memory | File-based memory enabled. Stored in ./.agent/memory |
| Skills | Loads Agent Skills from ./.agent/skills |
| Subagents | Delegates subtasks to the built-in generalist subagent |
| ToDo | Tracks multi-step tasks via checklists |
The programmatic_tool_caller allows the agent to call tools from within model-generated code.
The blog post explains that most of the token efficiency and benchmark accuracy advantages stem from these context management defaults.
Excessively long tool outputs are truncated and saved to disk, while conversation history is compressed before context limits are reached.
Swapping LLMs with a Single String
Models are specified as a provider/model-name string.
According to the model configuration guide, seven provider prefixes are supported: bedrock, bedrock-mantle, anthropic, openai, google, ollama, and litellm.
Strings without a prefix are treated as Amazon Bedrock model IDs.
# pip install strands-harness
from strands_harness import create_harness
agent = create_harness(model="openai/gpt-5.6-sol")
agent("Research the top three vector databases, compare pricing and limits, and write it up in comparison.md")
In TypeScript, this is written as createHarness({ model: "openai/gpt-5.6-sol" }).
Replacing the string with anthropic/claude-opus-5 or google/gemini-3.5-flash runs the same agent on another model, while local Ollama instances can be targeted using ollama/llama3.1.
Routing through LiteLLM enables connections to providers without native aliases.
Reasoning intensity is adjusted via effort with "low", "medium", "high", or "off".
Differences in reasoning parameters across providers are normalized by the library.
Specifying a level unsupported by a provider raises an error at harness creation time rather than mid-request.
Feature availability varies depending on the chosen model.
Native web_search is limited to providers including OpenAI, Anthropic, and Google. Prompt caching is explicitly managed by Strands harness only for Bedrock and direct Anthropic connections; OpenAI, Gemini, Bedrock Mantle, and LiteLLM handle caching automatically on the provider side, while Ollama does not support prompt caching.
Benchmark Numbers from Official Data
The official blog post presents two charts: one averaging six benchmark suites, and another focusing specifically on Terminal-Bench 2.1 for terminal-based tasks.
Both tables below are transcribed directly from raw numbers inside the embedded chart HTML.
AWS reports that measurements were run distributed across Amazon EC2 using the Harbor evaluation framework.
Average Across Six Benchmarks
The benchmark suite includes ALFWorld, ContextBench, GAIA, WebShop, τ³-bench, and Terminal-Bench 2.1.
Scores reflect average accuracy across all six, while costs denote average dollar spend per task.
The five competitors tested were Claude Code, Codex, oh-my-pi, OpenCode, and DeepSeek Harness.
The table below pairs Strands harness against Claude Code and Codex when configured with identical models:
| Model | Competitor | Competitor Score | Competitor Cost | Strands Score | Strands Cost |
|---|---|---|---|---|---|
| Claude Fable 5 | Claude Code | 80.94% | $1.872 | 81.11% | $0.829 |
| Claude Opus 5 | Claude Code | 84.60% | $1.063 | 84.55% | $0.634 |
| Claude Opus 4.8 | Claude Code | 77.32% | $0.983 | 77.14% | $0.421 |
| Claude Sonnet 5 | Claude Code | 71.55% | $0.974 | 75.61% | $0.311 |
| GPT-5.6 Sol | Codex | 76.59% | $0.628 | 76.72% | $0.293 |
| GPT-5.6 Terra | Codex | 59.06% | $0.178 | 64.75% | $0.143 |
| GPT-5.6 Luna | Codex | 62.66% | $0.028 | 61.94% | $0.009 |
For five of the seven pairs, score differences stayed under one percentage point.
The only notable gaps appeared with Claude Sonnet 5 (~4 points) and GPT-5.6 Terra (~6 points), where Strands harness scored higher in both cases.
Across all seven pairings, Strands harness cost less, with expenses reduced to between roughly 60% and 33% of Claude Code.
The headline “28%” refers to the combined average across all five competitors.
Chart footnotes indicate that DeepSeek Harness was ~14% cheaper than Strands harness, but lagged behind on all benchmark scores.
Including DeepSeek Harness in the overall comparison pulled the aggregated savings figure down to 28%.
Compared strictly against Claude Code, the cost reduction is substantially wider.
Terminal-Bench 2.1 with Claude Fable 5
The second chart compares all harnesses running Claude Fable 5 across 89 trials on Terminal-Bench 2.1:
| Harness | Accuracy | Cost | Tokens |
|---|---|---|---|
| DeepSeek Harness | 59.55% | $40.30 | 21.44M |
| Strands harness | 69.66% | $56.29 | 28.73M |
| OpenCode | 66.29% | $73.42 | 37.11M |
| oh-my-pi | 69.66% | $86.83 | 45.06M |
| Claude Code | 61.80% | $248.05 | 54.99M |
The “77% cheaper than Claude Code” claim comes directly from this dataset.
$56.29 is roughly 23% of $248.05, while Strands harness achieved a 7.86-point lead in accuracy.
While oh-my-pi tied Strands in accuracy at 69.66%, its total cost was about 1.5x higher.
Tokens differed by roughly 1.9x, yet cost diverged by approximately 4.4x.
The exact drivers behind this pricing gap are not detailed in the official blog post.
Caveats When Reading the Benchmark Data
These benchmarks were measured and published directly by AWS, with an academic paper promised for a later date.
The Register noted that while Strands harness is marketed as a general-purpose agent, all of its evaluation targets are dedicated coding agents.
The chart source code also disclosed several caveats regarding data collection:
| Note in Source | Practical Impact |
|---|---|
| Only pairs with complete cost data across all 6 benchmarks were plotted | Excluded eve, deepagents, and GPT-6 Astra. Astra was missing cost numbers for most benchmarks |
Missing cells were backfilled from a separate report (harbor-report-sep17) | Applied to Claude Code + Fable 5 on Terminal-Bench, 4 ContextBench points for oh-my-pi, and Codex + GPT-5.6 Terra on ALFWorld |
| Costs used measured values when available, falling back to estimates | The chart does not indicate which individual cells rely on estimates |
The Claude Code + Fable 5 row in the six-benchmark average borrows its Terminal-Bench datapoint from a separate run.
While the numbers above reflect published figures, these backfills and estimates provide essential context when evaluating the results.
Default Permissions and Unprompted Shell Execution
Coming from Claude Code, the most jarring default is tool execution permissions.
By default, Strands harness configures no intervention hooks (interventions), so every tool call executes immediately without confirmation.
Because shell and file operations execute directly on the local host environment, an agent initialized without arguments runs commands on your system unprompted.
According to the interventions documentation, four approval modes are available:
| Setting | Behavior |
|---|---|
interventions="ask" | Prompts for confirmation before every tool call |
interventions="smart" | Prompts only when the SDK’s risk classifier flags an action as hazardous |
| Natural language string | Uses guidelines like “confirm before file deletion or network requests” as criteria |
Path ending in .cedar | Enforces permissions defined in a Cedar policy file |
The built-in generalist subagent inherits these settings, so delegated subtasks cannot bypass approval gates.
The production deployment guide also emphasizes locking down approvals for shell and file edits alongside programmatic_tool_caller.
While programmatic_tool_caller executes model-authored code inside an isolated Monty environment, tools invoked from that code are not isolated.
For untrusted inputs, AWS recommends executing inside Docker or SSH sandboxes or disabling the tool altogether.
The shell tool spawns a fresh process for each invocation. Working directories and exported environment variables do not persist across calls.
File tools require absolute paths and reject paths containing ...
Deployment Targets
The library can be deployed anywhere Linux containers run. The blog highlights targets including Modal, Cloudflare Containers, Azure Container Apps, Google Cloud Run, Amazon ECS, and Amazon Bedrock AgentCore.
Because create_harness() returns a standard Strands Agent, deployment workflows written for the SDK apply directly.
When deploying to containers or serverless environments, storage configuration is critical.
Sessions and long-term memory default to ./.agent, which vanishes when ephemeral instances shut down.
Point them to persistent storage using session={"dir": ...} and memory={"dir": ...}.
Installation and the Strands CLI
The library can be installed via pip install strands-harness (Python 3.10+) or npm install @strands-agents/harness.
As of September 29, 2026, releases are 0.1.2 on PyPI and 0.1.1 on npm.
For interactive prototyping, the team provides the Strands CLI.
In the official demo, the CLI was used to select a model and tools, prompt the agent to “add Playwright MCP and measure video loading latency on a blog post,” and test execution.
Running /export afterward generates ready-to-use Python or TypeScript code incorporating the Playwright MCP configuration.
The CLI itself is built directly on top of Strands harness.