Gemini 4 Argon vs GPT-6.1 Sol & Opus 5.5: Pricing & CLI Tests on M4 Mac
Contents
Gemini 4 Argon was not yet available in Antigravity (AGY) on my M4 Mac mini. However, since GPT-6.1 Sol and Claude Opus 5.5 were accessible via their respective CLIs, I checked the official benchmark evals and ran small test prompts covering code generation and document fact-checking.
Both Sol and Opus offer performance close to each provider’s top-tier model at significantly lower costs. Gemini 4 Argon is Google’s latest frontier model, but as of October 1, access remains limited.
Availability in AGY
On October 1, 2026, I ran agy models on an M4 Mac mini to verify availability.
The latest Gemini models listed were Gemini 3.8 Flash (High, Medium, Low) alongside 3.1 Pro; no model names referencing 4 or Argon were present.
gemini-3.8-flash-high Gemini 3.8 Flash (High)
gemini-3.8-flash-medium Gemini 3.8 Flash (Medium)
gemini-3.8-flash-low Gemini 3.8 Flash (Low)
Google announced Gemini 4 Argon on September 30. At launch, availability is restricted to selected cybersecurity defenders participating in the Fairwind Program. Broader access is scheduled to roll out gradually starting with paid API users and Google AI Ultra subscribers.
Because Argon could not be called directly in this setup, I added Gemini 3.8 Flash High via AGY to the CLI test runs as a reference point.
Official Tiering and Price Brackets
OpenAI’s model selection guide categorizes Astra as peak intelligence, Sol as a balanced tier across speed, price, and capability, and Luna as high-speed and budget-friendly. Anthropic’s model overview structures its lineup into Fable, Opus, Sonnet, and Haiku, recommending Opus for daily engineering tasks and Fable for demanding multi-step reasoning.
| Model | Launch Date | Official Positioning |
|---|---|---|
| Gemini 4 Argon | Sept 30 | Frontier flagship; expanding from limited preview |
| GPT-6.1 Sol | Sept 29 | Astra-level performance at lower unit cost |
| Claude Opus 5.5 | Sept 22 | Matches Fable 5.1 on most engineering workflows |
Sol’s rollout is documented in the Codex and ChatGPT Work changelog, while Opus 5.5’s release and its comparison to Fable appear in Anthropic’s announcement. Sol is clearly a mid-tier refresh. In contrast, while Opus 5.5 is cheaper than the top-tier Fable, Sonnet and Haiku sit below it, making Opus more of a primary flagship workhorse than a traditional middle model.
API Pricing and Output Limits
In API billing, input tokens represent text fed into the model, while output tokens measure generated text. The table below lists regular API pricing per 1M tokens in US dollars, separate from flat subscription plans or overall session fees.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Max Output Tokens |
|---|---|---|---|
| Gemini 4 Argon (Introductory) | $2 | $10 | 1,000,000 tokens |
| Gemini 4 Argon (Post-Promo) | $4 | $20 | 1,000,000 tokens |
| GPT-6.1 Sol | $2 | $10 | 128,000 tokens |
| Claude Opus 5.5 | $4 | $20 | 128,000 tokens |
Argon’s numbers come from Google’s announcement, while Sol and Opus reflect their official model spec sheets (Sol spec, Opus spec). Sol’s listed rate applies to standard inputs up to 272,000 tokens; pricing shifts for longer context windows or high-speed routing modes.
Argon’s 1M token figure is the maximum generation length per response, not the input context window. Google expanded this limit from the previous 64,000 tokens to accommodate extended multi-turn reasoning and lengthy synthesis. For Opus, 128,000 tokens is the standard single-request cap, with a 300,000-token batch output option available in beta.
Opus 5.5’s token pricing dropped 20% compared to Opus 5. Anthropic’s claim of “40% lower cost per task” combines this rate reduction with fewer tokens required per task under default system settings.
Published Benchmark Comparisons
From Google’s published comparison table, the items directly contrasting Argon against Opus 5.5 are summarized below. OpenAI’s Sol 6.1 is not included in this official table.
| Benchmark | Task Description | Gemini 4 Argon | Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 | Long-horizon software engineering | 77.9% | 74.2% |
| Terminal-Bench 4.0 | Multi-step terminal execution | 57.4% | 66.4% |
| Vals Finance Agent v2 | Multi-turn financial research | 65.4% | 58.6% |
| LVBench | Long video comprehension | 91.7% | 83.7% |
Argon leads on DeepSWE and financial investigation, whereas Opus scores higher on Terminal-Bench. However, Google’s evaluation methodology doc indicates that Google-measured results and vendor self-reported figures are mixed here. Both use high reasoning effort by default, but the evaluation harnesses differ.
Video evaluation (LVBench) also involves different input conditions: Gemini evaluates at 1 fps, whereas Opus is constrained by API limits to a fixed 600 frames. These figures should not be read as a drop-in model swap within an identical environment.
While OpenAI characterizes Sol 6.1 as having “Astra-like intelligence,” the model spec page does not list scores on these specific benchmarks. Substituting Astra’s figures for Sol would compare different models.
Small CLI Benchmark Tests
To inspect real-world execution behavior, I sent identical benchmark prompts to each model from terminal CLIs.
Test Environment
| Item | Details / Version |
|---|---|
| Machine | M4 Mac mini (16 GiB unified memory / macOS 26.5.2) |
| GPT-6.1 Sol | Codex CLI 0.159.2 (Effort: medium) |
| Claude Opus 5.5 | Claude Code 2.1.285 (Effort: medium) |
| Gemini 3.8 Flash High (Reference) | AGY 1.2.14 (Effort: High) |
Each CLI was launched inside an empty scratch directory with web search and tool execution disabled, prompting the model to return code and answers in structured JSON. The prompts and grading harnesses are archived in the LiltingChannelLabo experiment repo.
| Task | Evaluation Criteria |
|---|---|
| Python interval merge function | Merge overlapping intervals, keep adjacent boundary points separate. 24 test cases covering edge cases, empty inputs, and large integers |
| Event aggregation | Parse 9 log lines, select the latest record per ID, and sum paid amounts grouped by currency |
| Document assertion check | Read a mock product spec and correctly categorize 3 claims distinguishing input/output, preview/GA, and pricing tiers |
Benchmark Results and Large Integer Handling
Running each prompt twice across the models produced the following results. Because event aggregation (2/2 correct runs) and document verification (3/3 questions correct) passed cleanly on all three models, the table isolates code test pass rates and elapsed wall-clock times.
| Model | Code Tests Passed | Wall-Clock Time (Run 1 → Run 2) |
|---|---|---|
| GPT-6.1 Sol (medium) | 24/24 → 24/24 | 20.8s → 18.3s |
| Claude Opus 5.5 (medium) | 24/24 → 24/24 | 18.5s → 19.8s |
| Gemini 3.8 Flash High (Reference) | 23/24 → 23/24 | 62.4s → 99.3s |
Sol and Opus matched each other across all functional tasks. While two trials per model are too few for rigorous latency benchmarking, response times hovered consistently around 19–20 seconds.
The single divergence occurred in the Python interval merge problem when testing with a large integer (10**400), where Gemini 3.8’s generated script crashed with OverflowError on both attempts.
Inspecting the code revealed that Gemini validated endpoint coordinates using math.isfinite(). While Python handles arbitrary-precision integers of arbitrary size, math.isfinite() implicitly casts inputs to a floating-point number (float). Because 10**400 exceeds standard IEEE 754 float limits (around 1.8 × 10³⁰⁸), the cast raised an overflow error.
# Generated by Gemini 3.8 Flash High
if not (math.isfinite(start) and math.isfinite(end)):
raise ValueError("endpoints must be finite numbers")
Sol and Opus avoided the crash by checking isinstance(x, float) before calling math.isfinite(), or by omitting the float-only validation for integers. Normal numeric ranges and boundary separation logic passed without issue across all three engines.