OpenAI’s GPT-6.1 Sol arrives with an unusually clear commercial proposition: capability close to the company’s premium Astra model at a much lower published API price. That sounds like an easy purchasing decision. It is not.
The useful question for a business is not “Which model has the lowest token price?” It is “Which model completes our real task acceptably, at the lowest total cost and risk?” Those are different calculations—and the gap between them can decide whether an AI workflow saves money or merely produces a cheaper stream of mistakes.
What OpenAI actually announced
OpenAI released GPT-6.1 Sol on 29 September 2026 for complex coding, computer use and professional work. Its launch documentation lists standard API prices of $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. The equivalent published prices for GPT-6 Astra are $10 input, $1 cached input and $50 output.
OpenAI reports that Sol approached Astra on several internal or third-party evaluations while costing substantially less per task. It matched Astra on the company’s reported DeepSWE software-engineering result at roughly one-fifth of the cost, came within 2.1 percentage points on an OSWorld computer-use evaluation at roughly one-seventh of the cost, and remained below Astra on the hardest scientific-workflow evaluation.
Those results are promising, but they are not a universal purchasing verdict. OpenAI explicitly notes that its evaluations ran in research or API environments that may differ from production. Benchmark prompts, tools, reasoning settings and failure costs may not resemble yours.
Why token price is only the first line of the bill
A model invoice captures tokens and tool calls. A business case must also capture everything required to turn an answer into an acceptable outcome.
- Retries: How often must the workflow run again?
- Human review: How many minutes does checking or correcting each result take?
- Latency: Does waiting create abandoned sessions, slower support or idle staff?
- Tool costs: Are search, computer use, storage or third-party services charged separately?
- Failure impact: What does an incorrect refund, broken code change or unsupported claim cost?
- Integration: What engineering, monitoring and migration work does the model require?
The current GPT-6.1 Sol API documentation also shows why procurement teams should read beyond the headline rate. Prompts above 272,000 input tokens attract higher rates for the full request; Fast mode costs twice the standard rate, Ultrafast costs six times standard, and regional processing can add a 10% premium where available. Tool-specific charges may apply too.
Use cost per successful task
A more useful metric is:
Cost per successful task = total workflow cost ÷ number of outputs that pass the acceptance test
“Successful” must be defined before the test. For a support-drafting workflow, success might mean a correct answer grounded in approved policy, the right tone, no prohibited disclosure and no material edit by the reviewer. For code, it might mean tests pass, security checks pass and a maintainer accepts the change.
An illustrative calculation
Suppose a workflow averages 6,000 uncached input tokens and 2,000 output tokens. At the published standard rates, the model-token cost would be approximately $0.032 per GPT-6.1 Sol attempt and $0.16 per Astra attempt. Across 10,000 attempts, that is about $320 versus $1,600 before caching, tools, retries, review and other charges.
Now add a hypothetical acceptance test. If Sol passes 82% of tasks and Astra passes 90%, model cost per accepted result is about $0.039 for Sol and $0.178 for Astra. Sol still looks attractive. But if every failed result takes a skilled employee ten minutes to repair, human correction cost can dwarf both model bills. A small accuracy gap may matter greatly in an expensive or high-risk workflow.
These success rates are illustrative, not measured claims about either model. Your own evaluation must supply them.
A practical model-selection test
1. Choose one bounded workflow
Do not begin with “Which model should our company use?” Begin with one decision such as triaging support tickets, extracting clauses from contracts or preparing a first-pass product description. Different tasks may justify different models.
2. Build a representative test set
Include routine cases, ambiguous cases, long inputs, missing information and examples where the correct action is to stop or escalate. Remove personal or confidential data unless your approved environment and governance permit its use.
3. Write the acceptance test first
Score dimensions that matter to the business: factual correctness, completeness, policy compliance, format, latency and reviewer effort. Weight serious failures more heavily than cosmetic ones. A fluent but unauthorised action is not “almost correct”.
4. Hold the workflow constant
Use the same prompts, tools, context and output requirements for each candidate where possible. Record reasoning effort, caching and service tier. Otherwise, you may compare configurations rather than models.
5. Measure the whole system
For every run, capture token and tool cost, elapsed time, pass or fail, retry count and human-review minutes. Then calculate cost per accepted result—not just cost per request.
6. Run a guarded pilot
Start in assistive mode. Let the system recommend, draft or classify while a person retains authority. Increase autonomy only when repeated evidence supports it and a rollback path exists.
The safety card changes the procurement conversation
GPT-6.1 Sol is not merely a cheaper text generator. OpenAI’s system-card addendum says the company treats it as reaching its Critical capability threshold in cybersecurity and High in biological and chemical capability, applying the same safeguards stack used for Astra.
That does not mean ordinary business use is inherently unsafe. It means stronger capability deserves deliberate access controls, monitoring and escalation. The lowest-cost capable model may increase the number of workflows an organisation can afford to automate; it also increases the number of places where governance can fail.
The NIST AI Risk Management Framework offers a useful vendor-neutral discipline: map the context and impacts, measure performance and risk, manage identified problems, and govern the system throughout its lifecycle. NIST’s generative-AI profile specifically recommends documenting context of use, foreseeable misuse, measurement plans and human-AI configurations.
Who should consider GPT-6.1 Sol?
It is a plausible candidate for teams that need stronger coding, document analysis, computer-use or multi-step workflow performance than a low-cost everyday model provides, but cannot justify premium-model pricing on every request. It may also suit routing designs in which a lower-cost model handles most cases and difficult or high-risk cases escalate to Astra or a human expert.
It is a weaker fit when a lightweight model already meets the acceptance threshold, when latency or on-device operation matters more than reasoning depth, or when the organisation lacks the evaluation and oversight needed for agentic work. Doing nothing—or improving the existing non-AI workflow—remains a valid alternative when errors are costly and task volume is low.
The strategic signal behind the release
The important trend is not one model winning a benchmark. It is frontier-level capability moving down the cost curve while model portfolios become more specialised by price, speed and risk. That makes “use the most capable model everywhere” increasingly hard to defend.
A mature AI stack will probably look less like a single subscription and more like a routing policy: small models for predictable high-volume work, stronger models for ambiguity, premium models for the hardest cases, and people for accountability and exceptions.
The takeaway
GPT-6.1 Sol’s published price is significant, and OpenAI’s reported evaluations make it worthy of testing. Neither proves that it is the cheapest model for your business.
Run the same representative work through realistic configurations. Count accepted outcomes, review time, retries, tool charges and failure impact. Then choose the model—or combination of models—that delivers the best verified result for the workflow.
Featured image: conceptual AI illustration created for MaryChuks.com; not documentary evidence.
Discover more from Marychuks.com AI, Psychology, Business & CreativeVerse
Subscribe to get the latest posts sent to your email.