I would rather review two actual outputs than spend an afternoon arguing about model rankings. The best engine on paper can still produce a poor account brief when it receives stale notes and a vague definition of finished.
My rule is to buy the next increment of reasoning only after naming the failure it should fix. I call that an earned upgrade. The price of the experiment belongs next to the quality of the result.
Separate the three choices
First choose the working environment: Chat, Work, Codex or your API integration. Then choose an available model. Finally decide how much reasoning effort the task deserves. Those choices interact, but collapsing them into “which ChatGPT is best?” hides the actual decision.
For a GTM operator, the environment determines whether the task can reach the account evidence and whether you can review the output in the form you need. The model determines how well it reasons over that evidence. Effort influences how much work it puts into the answer. A different model cannot compensate for a CRM connector that has never been authenticated.
Write the three choices in your experiment log. Otherwise a faster result could reflect a changed tool setup, shorter source packet or different effort setting rather than a model improvement. Keep one variable moving at a time.
TipRecord product, model and effort together.
Read the current model guidance in scope
As checked September 22, 2026, official Work and Codex guidance positions GPT-6 Astra for demanding work across steps and tools, Sol for complex or open-ended work, Terra for everyday tasks and Luna for clear, repeatable work. These are useful starting hypotheses, not a promise that every account has the same picker.
Astra rollout, plan access and organization settings can limit availability. The controls beneath the composer may expose a Power slider and an Advanced option rather than the exact list shown in someone else's tutorial. For Codex CLI, /model is the place to inspect selectable options and reasoning effort in an interactive session.
This guide does not claim that the Work and Codex menu is the universal Chat menu or API model catalog. Check the specific surface. I would rather show that boundary explicitly than publish a single tidy table that sends half the readers looking for an option they cannot access.
TipTreat your actual model picker and account access as the operational source of truth.
Match capability to the shape of the job
A short classification task has a narrow output and an explicit rubric. That is a good candidate for evaluating a faster option such as Luna. A campaign analysis with several inconsistent sources asks for more judgment. A code change touching multiple parts of a repository asks for sustained reasoning and verification.
For fictional Cedar Metrics, classify inbound records into fit, not fit or insufficient information. Compare those labels against a manually reviewed sample. Then use a separate test for the weekly pipeline narrative, where the harder question is whether the evidence supports the explanation. A model can do well on one and poorly on the other.
I would avoid routing by department alone. “Marketing uses Terra” sounds administratively neat but says little about task difficulty. Routing by the work is more useful: extraction, synthesis, complex reasoning and multi-step execution. The department is who owns the workflow; it is not the evaluation.
TipGive ambiguous and incomplete records their own cases in the test set.
Adjust effort before making a grand conclusion
Reasoning effort changes the resources devoted to a task. Current interfaces can show labels such as Light, Medium, High and Extra High; CLI terminology can differ. Higher settings can take longer and consume more usage. Compare the completed result, not the amount of visible activity.
Suppose the pipeline narrative fails to reconcile a changed close date with a more recent call note. Repeat the same task with stronger reasoning while keeping the evidence fixed. If the error remains, inspect whether the source packet itself identifies which date is authoritative. More computation over an unresolved business rule can produce a longer unresolved answer.
Max and Ultra deserve separate attention. Current documentation describes Max as additional reasoning time and Ultra as using subagents for separable work. Those are not interchangeable speed buttons. Use parallel work when the task can be divided into meaningful independent parts and you have a way to reconcile their results.
TipUse the lowest setting that meets the acceptance criteria, then retest when the task changes.
Build a small evaluation you can explain
An evaluation is a set of inputs with a clear definition of acceptable output. For lead classification, collect representative records across segments, including missing employee counts and contradictory descriptions. Mark the expected classification and why. Keep a few examples out of the prompt so you can test whether the workflow generalizes.
For a meeting brief, score unsupported claims, missing material facts, stale dates and editing time. Avoid judging the whole answer with one impressionistic star rating. A beautiful brief that invents stakeholder support should fail a meaningful check, even if it wins on tone.
Use the same inputs across candidate settings. Save the actual output, source packet, model and date. If someone asks why the team selected a particular option, you can show evidence. That beats “it felt smarter on Tuesday,” which is a wonderfully human procurement method and a poor one.
TipKeep quality criteria visible before comparing outputs.
Measure cost per accepted result
A cheap request that needs three retries and ten minutes of editing may cost more than a stronger first result. Count the whole workflow: model usage, tools, retries, review and the operational cost of errors. App allowances and API token prices are different billing systems, so compare within the route you actually use.
Here is an illustrative calculation, not a benchmark: ten briefs cost a hypothetical $4 in metered usage and require 30 minutes of review at an assumed $40 hourly internal rate. That is $24 total, or $2.40 per accepted brief if all ten pass. If only eight pass, the denominator changes. The assumptions should be visible so finance can replace them.
You do not need perfect accounting to learn something. A simple ledger of accepted outputs and review minutes often reveals whether the bottleneck is model quality, missing sources or an overlong format. Start there before redesigning the stack.
TipUse an explicitly assumed labor rate; do not present a hypothetical saving as a measured result.
Handle retirements as configuration work
The official documentation currently says GPT-5.5 retires from ChatGPT, Work and Codex on October 14, 2026, while the API is outside that notice. For Codex with ChatGPT sign-in, it names GPT-5.6 Sol as the replacement. This is a dated notice to verify again before making changes.
Search saved defaults, scheduled tasks, custom agent configurations and scripts for an explicit old model. Changing the picker in one conversation does not update every workflow. Confirm access to the replacement, replay representative tests and record any behavior differences before moving production work.
Preserve a working version of prompts and settings so the team can explain the change. If a model has actually retired, rollback may require a different supported option rather than the unavailable model. Operational planning starts with what can still run, not with a screenshot of last month's successful task.
TipKeep retirement dates in a maintenance log with an owner.
Where model selection goes wrong
The biggest trap is treating a missing feature as weak intelligence. No model can inspect an unavailable spreadsheet. Another is upgrading from one attractive demonstration. Your real workload includes sparse sources, conflicting fields and dull repetitive tasks that rarely appear in launch videos.
A third trap is confusing plan access with runtime permissions. Selecting a model does not grant it access to a folder or external service. Likewise, granting broader file access does not make an unavailable model selectable. The permissions guide separates those controls.
My practical recommendation is to keep a default, a harder-task option and a few representative tests. Revisit the choice when the work or product changes. Which failure in your current output would a stronger model actually need to solve?
How to set it up
Define acceptance
Choose one workflow and write its hard failures: unsupported account facts, invalid labels, incorrect totals or broken tests.
Record the baseline
Use an available default with the same source packet and output instructions. Record the model, effort, product and date.
Compare one alternative
Change the model or effort, not both at once. Inspect factual quality, review time and usage.
Update the workflow
Choose the lowest-cost setup that passes. Review saved model references and document the decision with the test outputs.
Frequently asked questions
Which model is best?
It depends on the task and available surface. Use official positioning to shortlist options, then compare representative outputs.
Are Chat and Codex model menus identical?
Do not assume so. Availability can depend on product, plan, client, rollout and sign-in method.
What is reasoning effort?
A control over how much reasoning work a supported model devotes to the task, affecting time and usage.
Does Ultra mean a stronger single model?
Current documentation describes Ultra as using subagents for parallel parts of complex work. It differs from simply increasing single-task reasoning time.
Should every task use Astra?
Evaluate it where demanding reasoning or execution warrants the cost. Clear routine work deserves its own comparison.
Does the GPT-5.5 retirement include the API?
The current October 14, 2026 notice excludes the OpenAI API. Verify the notice and your authentication route.
Can a model read every connected source?
Only sources accessible through the task environment and authorized tools. Model selection does not grant access.
How often should I retest?
When the model, prompt, source structure or workflow changes, and when a material error exposes a gap in the existing tests.
Sources & further reading
- Current models and reasoning controls
- Product-specific model access
- Usage and pricing
- Work capabilities
ChatGPT and Codex change quickly. This page was last reviewed September 22, 2026; verify time-sensitive details against the official docs above before relying on them.