By Mel RossBook a call

OpenAI's new model costs a fifth as much. Test it on your own work.

October 7, 2026

I'd skip the benchmark scores in this week's OpenAI coverage. They may be accurate, and they still won't help you decide anything.

On September 30, OpenAI released a model called GPT-6.1 Sol. The pitch is that performance is close to OpenAI's top model, GPT-6 Astra, at roughly a fifth of Astra's price.

Then come the scores. Sol beat the older GPT-6 Sol by 6.4 points (on a test called DeepSWE v1.1), matching Astra's result there at a fifth of the cost. It beat a competitor's model, Claude’s Opus 5.5, by 2.2 points on AutomationBench at a third of the cost. It beat GPT-6 Sol by 7 points on OSWorld 2.0 at half the cost.

If you run a landscaping company or a 6-person accounting office, none of those names mean anything to you. That's fine. Benchmarks are scored tests AI companies run to compare models, and the coverage describes these ones as tests of "agentic, coding, and professional tasks." Agentic means the AI carries out steps on its own, like clicking through software, instead of just answering you.

Neither article says how a 6.4-point gain shows up when you ask for a reply to a customer who's furious about a late delivery. Or when you need a contract clause read for the one line that matters. Or a spreadsheet formula that actually adds up the right column.

That gap is the whole problem with launch coverage. A company ships a model, the write-ups lead with scores, and you're left guessing whether to switch. Then somebody else ships one with better scores, and you're guessing again.

Who can actually use it

This is the detail most coverage moves past fast. Sol isn't on the free ChatGPT tier, and it isn't in the regular ChatGPT chat window. It's in ChatGPT Work and Codex for people on Plus, Pro, Business, Enterprise and Edu plans. Developers can also reach it through the API, which is the connection software companies use to build OpenAI's models into their own products.

So if you open ChatGPT on a free account this afternoon to try it, you won't find it. If you pay for Plus, look in ChatGPT Work. Your usual chat window doesn't have it.

The price is listed per token: $2 per million input tokens, $10 per million output tokens, and $0.10 per million for cached input. A token is essentially a chunk of a word. Input is what you send; output is what it writes back. If you pay a flat monthly fee for ChatGPT, those prices aren't on your bill. They matter to the software companies building on OpenAI, which now have a much cheaper option close to the top model. Whether any of that savings reaches the tools you pay for is up to them.

Build your own test

Here's my position. The only test that tells you whether an AI model is good enough for your business is one made of your business's work. Build it once, keep it, and rerun it every time a launch makes you wonder.

  1. Pull 3 real tasks from last week. Pick things you'd actually hand to AI: a customer email that needed care, or a formula you had to look up. Make sure at least 1 has an answer you already know is right.
  2. Write the exact prompt for each and save it in a doc. Same words every time, or the comparison means nothing. Strip out client names and anything private first.
  3. Run all 3 through whatever you use now. Paste each answer under its prompt with the date and the model name.
  4. When the next launch shows up, run the same 3 prompts on the new model. Put the answers side by side and read them. Would you send that email? Is the number right?

The first round is the slow one. Every round after that is 3 copy-pastes and some reading.

The task with a known answer is the one that earns its keep. A nicely written wrong answer looks fine unless you already know what right looks like. A summary that sounds confident and skips the clause that matters will fool you every time if you're grading on tone.

This also fixes the "which one is better" argument in your office. Somebody on your team swears by one tool, somebody else swears by another. Run the same 3 prompts through both and read the results together. Now the argument is about answers to your own work.

So here's the decision for this week: pick your 3 tasks and save the prompts before you touch any settings or upgrade anything. If Sol is in a plan you already pay for, run your test on it. If it isn't, run the test on what you have now. You'll have a baseline ready for the next launch, and there'll be one.

If you'd like your whole team building tests like this from their own work, and knowing how to read the results, book a call and I'll train them on it.

Source: TestingCatalog, "OpenAI launches GPT-6.1 Sol at one-fifth of Astra pricing", September 30, 2026; confirmed by APH Networks.

More like this, weekly.

One email a week on operations, systems and where AI actually helps.