Tech6 min read

Gemini 4 Argon vs GPT-6.1 Sol & Opus 5.5: Pricing & CLI Tests on M4 Mac

IkesanContents

Gemini 4 Argon was not yet available in Antigravity (AGY) on my M4 Mac mini. However, since GPT-6.1 Sol and Claude Opus 5.5 were accessible via their respective CLIs, I checked the official benchmark evals and ran small test prompts covering code generation and document fact-checking.

Both Sol and Opus offer performance close to each provider’s top-tier model at significantly lower costs. Gemini 4 Argon is Google’s latest frontier model, but as of October 1, access remains limited.

Availability in AGY

On October 1, 2026, I ran agy models on an M4 Mac mini to verify availability. The latest Gemini models listed were Gemini 3.8 Flash (High, Medium, Low) alongside 3.1 Pro; no model names referencing 4 or Argon were present.

gemini-3.8-flash-high    Gemini 3.8 Flash (High)
gemini-3.8-flash-medium  Gemini 3.8 Flash (Medium)
gemini-3.8-flash-low     Gemini 3.8 Flash (Low)

Google announced Gemini 4 Argon on September 30. At launch, availability is restricted to selected cybersecurity defenders participating in the Fairwind Program. Broader access is scheduled to roll out gradually starting with paid API users and Google AI Ultra subscribers.

Because Argon could not be called directly in this setup, I added Gemini 3.8 Flash High via AGY to the CLI test runs as a reference point.

Official Tiering and Price Brackets

OpenAI’s model selection guide categorizes Astra as peak intelligence, Sol as a balanced tier across speed, price, and capability, and Luna as high-speed and budget-friendly. Anthropic’s model overview structures its lineup into Fable, Opus, Sonnet, and Haiku, recommending Opus for daily engineering tasks and Fable for demanding multi-step reasoning.

ModelLaunch DateOfficial Positioning
Gemini 4 ArgonSept 30Frontier flagship; expanding from limited preview
GPT-6.1 SolSept 29Astra-level performance at lower unit cost
Claude Opus 5.5Sept 22Matches Fable 5.1 on most engineering workflows

Sol’s rollout is documented in the Codex and ChatGPT Work changelog, while Opus 5.5’s release and its comparison to Fable appear in Anthropic’s announcement. Sol is clearly a mid-tier refresh. In contrast, while Opus 5.5 is cheaper than the top-tier Fable, Sonnet and Haiku sit below it, making Opus more of a primary flagship workhorse than a traditional middle model.

API Pricing and Output Limits

In API billing, input tokens represent text fed into the model, while output tokens measure generated text. The table below lists regular API pricing per 1M tokens in US dollars, separate from flat subscription plans or overall session fees.

ModelInput (per 1M tokens)Output (per 1M tokens)Max Output Tokens
Gemini 4 Argon (Introductory)$2$101,000,000 tokens
Gemini 4 Argon (Post-Promo)$4$201,000,000 tokens
GPT-6.1 Sol$2$10128,000 tokens
Claude Opus 5.5$4$20128,000 tokens

Argon’s numbers come from Google’s announcement, while Sol and Opus reflect their official model spec sheets (Sol spec, Opus spec). Sol’s listed rate applies to standard inputs up to 272,000 tokens; pricing shifts for longer context windows or high-speed routing modes.

Argon’s 1M token figure is the maximum generation length per response, not the input context window. Google expanded this limit from the previous 64,000 tokens to accommodate extended multi-turn reasoning and lengthy synthesis. For Opus, 128,000 tokens is the standard single-request cap, with a 300,000-token batch output option available in beta.

Opus 5.5’s token pricing dropped 20% compared to Opus 5. Anthropic’s claim of “40% lower cost per task” combines this rate reduction with fewer tokens required per task under default system settings.

Published Benchmark Comparisons

From Google’s published comparison table, the items directly contrasting Argon against Opus 5.5 are summarized below. OpenAI’s Sol 6.1 is not included in this official table.

BenchmarkTask DescriptionGemini 4 ArgonOpus 5.5
DeepSWE v1.1Long-horizon software engineering77.9%74.2%
Terminal-Bench 4.0Multi-step terminal execution57.4%66.4%
Vals Finance Agent v2Multi-turn financial research65.4%58.6%
LVBenchLong video comprehension91.7%83.7%

Argon leads on DeepSWE and financial investigation, whereas Opus scores higher on Terminal-Bench. However, Google’s evaluation methodology doc indicates that Google-measured results and vendor self-reported figures are mixed here. Both use high reasoning effort by default, but the evaluation harnesses differ.

Video evaluation (LVBench) also involves different input conditions: Gemini evaluates at 1 fps, whereas Opus is constrained by API limits to a fixed 600 frames. These figures should not be read as a drop-in model swap within an identical environment.

While OpenAI characterizes Sol 6.1 as having “Astra-like intelligence,” the model spec page does not list scores on these specific benchmarks. Substituting Astra’s figures for Sol would compare different models.

Small CLI Benchmark Tests

To inspect real-world execution behavior, I sent identical benchmark prompts to each model from terminal CLIs.

Test Environment

ItemDetails / Version
MachineM4 Mac mini (16 GiB unified memory / macOS 26.5.2)
GPT-6.1 SolCodex CLI 0.159.2 (Effort: medium)
Claude Opus 5.5Claude Code 2.1.285 (Effort: medium)
Gemini 3.8 Flash High (Reference)AGY 1.2.14 (Effort: High)

Each CLI was launched inside an empty scratch directory with web search and tool execution disabled, prompting the model to return code and answers in structured JSON. The prompts and grading harnesses are archived in the LiltingChannelLabo experiment repo.

TaskEvaluation Criteria
Python interval merge functionMerge overlapping intervals, keep adjacent boundary points separate. 24 test cases covering edge cases, empty inputs, and large integers
Event aggregationParse 9 log lines, select the latest record per ID, and sum paid amounts grouped by currency
Document assertion checkRead a mock product spec and correctly categorize 3 claims distinguishing input/output, preview/GA, and pricing tiers

Benchmark Results and Large Integer Handling

Running each prompt twice across the models produced the following results. Because event aggregation (2/2 correct runs) and document verification (3/3 questions correct) passed cleanly on all three models, the table isolates code test pass rates and elapsed wall-clock times.

ModelCode Tests PassedWall-Clock Time (Run 1 → Run 2)
GPT-6.1 Sol (medium)24/24 → 24/2420.8s → 18.3s
Claude Opus 5.5 (medium)24/24 → 24/2418.5s → 19.8s
Gemini 3.8 Flash High (Reference)23/24 → 23/2462.4s → 99.3s

Sol and Opus matched each other across all functional tasks. While two trials per model are too few for rigorous latency benchmarking, response times hovered consistently around 19–20 seconds.

The single divergence occurred in the Python interval merge problem when testing with a large integer (10**400), where Gemini 3.8’s generated script crashed with OverflowError on both attempts. Inspecting the code revealed that Gemini validated endpoint coordinates using math.isfinite(). While Python handles arbitrary-precision integers of arbitrary size, math.isfinite() implicitly casts inputs to a floating-point number (float). Because 10**400 exceeds standard IEEE 754 float limits (around 1.8 × 10³⁰⁸), the cast raised an overflow error.

# Generated by Gemini 3.8 Flash High
if not (math.isfinite(start) and math.isfinite(end)):
    raise ValueError("endpoints must be finite numbers")

Sol and Opus avoided the crash by checking isinstance(x, float) before calling math.isfinite(), or by omitting the float-only validation for integers. Normal numeric ranges and boundary separation logic passed without issue across all three engines.