AI Model Selection: Choose Per Task, Not by Hype

Choosing an AI model is a workload decision, not a leaderboard decision. Use a scorecard for quality, latency, cost, reliability, and governance, then keep the architecture reversible.

Jason Bao
Abstract blue geometric field representing workload-based AI model selection

The short answer

AI model selection is the practice of matching a model to a specific workload by testing candidates on representative tasks and scoring them against requirements set in advance: output quality, latency, context size, tool-use reliability, data residency, and cost per completed task. Leaderboard rank is a screening filter. Your own test set makes the decision.

Mozilla's guide to choosing ML models shows why. Pegasus and BART both published ROUGE-1 summarization scores above 40. Run against Mozilla's own text, the same models scored roughly 27 to 31. Same models, same metric, different data. Microsoft's model selection guide is blunter still, naming two factors to deliberately exclude from the decision: cultural popularity, and who published the model.

What AI model selection actually is

Three decisions get collapsed into one, and they have very different lifespans. Model choice is which set of weights answers a request, and it turns over every few months on the provider's schedule rather than yours. Routing is the policy mapping a request to a model, and it changes when your traffic mix or cost target changes. Application architecture is where the workflow's logic, state, permissions, and audit trail live, and it should change only when the business process does.

The failure mode is letting the fastest-moving decision drive the slowest-moving one. When prompt text, retry behavior, output schemas, and business rules are scattered through code written against one provider's SDK, swapping a model stops being a config change and becomes a migration. Microsoft's derisking advice points the same way: put an abstraction layer between the application and the model, test candidates in parallel by switching environment variables, and avoid opaque routing unless you get observability with it.

The criteria that matter

Quality on your tasks

IBM defines model selection as choosing the candidate that generalizes best to unseen data, proven with holdout splits, k-fold cross-validation, and bootstrapping. You rarely retrain a foundation model, so the equivalent discipline is a frozen evaluation set you build once and reuse on every candidate.

Thirty to a hundred graded examples per task type separates real differences from noise on most enterprise workloads. Write the acceptance criteria before you look at any output, because grading after the fact is how a team talks itself into the model it already wanted. Benchmark rank is a screening signal. It tells you which four models are worth testing, not which one handles your contract redlines.

Latency and throughput

Measure p95. An agent that answers in 900 milliseconds most of the time and 14 seconds on the tail feels broken to the people using it, and the average hides that completely.

Separate time to first token from total completion time. A streaming chat interface lives on the first number, a batch job writing 4,000 records on the second. Reasoning models add compute beyond the visible tokens, which surfaces as latency you did not budget for. Rate limits belong in this row too. A model that clears your quality bar and then throttles at 40 requests per minute has failed a throughput requirement.

Context and tool use

An advertised context window is capacity, not usable capacity. Test retrieval and synthesis at 60 to 80 percent of the stated limit, with the relevant fact placed somewhere awkward.

Tool use needs its own measurements: malformed call rate, wrong-tool rate, and recovery behavior when a tool returns an error instead of a result. That last one separates models more sharply than any public leaderboard, and it decides whether an agent can run unattended.

Cost and token economics

Price the workload, not the model. Cost per million tokens is a rate card. Cost per completed task is what lands on your invoice, including retries, rejected outputs, and the tokens a re-prompt burns when schema validation fails.

Microsoft names three characteristics that move cost independently of volume: larger context windows raise input processing cost, multimodal inputs add tokenization overhead, and reasoning capability adds compute you are not billed for as visible output. A cheaper model that needs two attempts is not cheaper. Our guide to LLM cost optimization covers the unit economics. One structural point sits underneath the arithmetic. If every run of a repeatable task re-reasons from scratch, spend scales with usage forever, and every model swap re-prices the whole workload.

Privacy, residency, and governance

Region availability, data retention terms, and certification coverage decide eligibility before any quality score matters. Teams in healthcare, finance, and government carry GDPR, HIPAA, or CCPA obligations a strong eval result does not waive. IBM adds that interpretability carries weight in regulated fields, where an unexplainable decision is a compliance problem regardless of accuracy. Treat these as gates rather than weights. A model that fails a residency requirement does not get a score. It leaves the pool.

A weighted scorecard

The weights below are editorial defaults, sized for a tool-using enterprise agent. Retune them against your own workload before you trust the output. A real-time support surface pushes latency up and context down. A contract analysis pipeline does the reverse.

| Criterion | Weight | How to measure | Gate | |---|---|---|---| | Task quality on your eval set | 25% | Graded pass rate across five task types | Below 85% disqualifies | | Instruction and schema adherence | 15% | Valid structured output on first attempt | Below 95% disqualifies | | Tool-call reliability | 12% | Malformed and wrong-tool call rate | Above 3% disqualifies | | Cost per completed task | 10% | Total tokens including retries, at your mix | Budget ceiling | | Latency, p95 | 10% | Time to first token and total completion | Surface-specific ceiling | | Effective context handling | 8% | Retrieval plus synthesis at 70% of window | None | | Throughput and rate limits | 8% | Sustained requests per minute at peak | Peak load requirement | | Data residency and retention | 7% | Hosting region, retention terms, certifications | Hard gate, pass or fail | | Deployment options | 3% | Serverless, managed, self-hosted, on-device | Hard gate if on-prem required | | Change and deprecation policy | 2% | Notice period and version pinning support | None |

Run every candidate through the same five tasks, with identical prompts, identical data, and identical tool definitions, inside the same measurement window. Changing any of those between candidates invalidates the comparison.

  1. Structured extraction. Pull named fields from a messy real document, one with a scanned page and an inconsistent table.
  2. Code generation. Modify an actual file from your repository, not a puzzle, and check whether the result compiles and passes existing tests.
  3. Support triage. Classify and route a ticket against a written policy document, including the cases the policy does not cover.
  4. Long-context reasoning. Synthesize an answer across a set of documents where the relevant passages sit far apart.
  5. Multimodal input. Turn a screenshot or a scanned PDF into structured output. Multimodal means the model accepts more than one input type. Running more than one model is a separate question, covered below under routing.

The selection rule, stated plainly:

eligible = { m : m passes every hard gate }
score(m) = Σ wᵢ × normalizeᵢ(m) for all weighted criteria i
choose argmax score(m)
subject to p95_latency(m) ≤ L_max
and cost_per_task(m) ≤ C_max

Normalize each criterion to a 0-to-1 range across the candidate set before weighting, so a wide spread in one dimension does not swamp the rest.

Hold one separation while reading the results. This scorecard measures the model, and it says nothing about your system. Production reliability compounds model quality with retry policy, timeout handling, schema validation, tool availability, and the state your application keeps between steps. A model that passes at 94 percent inside a system with no retry logic and no durable state will not deliver 94 percent to a user. Reporting eval scores as service levels is a common way an AI program loses credibility with its own leadership.

When to use one model versus routing

Default to one model. A single default is easier to forecast and easier to debug, and Microsoft's guidance agrees that continuing with a proven model beats running a long evaluation you did not need.

Routing earns its complexity when request difficulty varies widely and the cheap requests dominate by volume. Cost-optimized routing sends the simple ones to a smaller model, quality-optimized routing pushes the hard or high-risk ones to a stronger one. Two constraints come with it. A router's effective context window is bounded by the smallest window in its backing pool, and runtime routing makes cost forecasting and debugging harder than a static choice. Our LLM router and LLM orchestration pieces cover the mechanics.

Configure fallback and escalation separately. Fallback handles provider failure: a timeout or a 503 sends the identical request, under the identical prompt contract, to a secondary model. Escalation handles difficulty: a low-confidence result, a high-value record, or a policy-flagged case goes to a stronger model or to a person.

Write escalation rules as explicit conditions in your application, tied to thresholds you can read and change, and log every escalation with the record it touched. A model deciding on its own that it feels unsure gives you no audit trail. A rule sending any refund above a stated amount to a named human queue gives you one.

The Major take

Model choice moves faster than the workflows it serves. Providers ship, deprecate, and reprice on a cadence measured in months, while a collections process stays recognizably the same for years. Every re-selection puts quality and spend back in play, and the more of your workflow lives inside prompts and model calls, the more you re-test each time a vendor moves.

The resolution is architectural. Enterprise AI agents holding an entire process in a context window re-reason it on every run, so a model change touches everything at once. On Major, when an agent works out how to handle a repeatable part of a task, it builds an app for that part and runs the app from then on. The extraction rules, the routing thresholds, the escalation queue, and the record of what happened live in deterministic code with a managed database, permissions, and logs behind them. The model keeps the judgment calls.

That makes selection reversible. Swapping the judgment model changes which weights answer a question. It does not change how the invoice gets reconciled, where the state lives, or who is allowed to see it. Your scorecard shrinks to the decisions the model genuinely owns, and the repeatable work stops being re-priced every time a provider ships. Reason once, run forever.

None of this removes the need to choose well. If your workload is a single stateless call with no repeatable structure underneath it, a scorecard and a decent abstraction layer will serve you. The case for a deterministic app layer gets stronger as the work gets more repetitive, more stateful, and more governed.

Pick your default model with the scorecard above, then put the repeatable half of the workflow somewhere a model swap cannot reach it. On Major, that means the agent builds the extraction, routing, and escalation steps into a governed app with its own database and audit log, and you keep testing models against the judgment that is left. Start building your model-agnostic agent on Major.

Related articles

Frequently asked questions

How do I choose an AI model?
Start with the workload, not a leaderboard. Write down your requirements for quality, latency, context, tool use, cost, and data residency, then build a frozen evaluation set of 30 to 100 graded examples per task type. Run every candidate on identical prompts, data, and tool definitions, score against weights you set in advance, and treat residency and deployment limits as pass-fail gates.
What factors matter most in AI model selection?
Quality on your own tasks carries the most weight, followed by instruction and schema adherence, tool-call reliability, cost per completed task, and p95 latency. Effective context handling and sustained throughput come next. Data residency, retention terms, and deployment options work as gates rather than scores, since a model that fails a compliance requirement is ineligible regardless of how well it performs.
Should an enterprise use one AI model or several?
Default to one. A single model is easier to forecast, debug, and govern. Routing earns its complexity when request difficulty varies widely and cheap requests dominate by volume, sending simple work to a smaller model and hard or high-risk work to a stronger one. Two costs come with it: a router's effective context window is bounded by the smallest window in its pool, and runtime routing complicates cost forecasting.
How do I compare AI models fairly?
Hold everything except the model constant. Use identical prompts, identical input data, identical tool definitions, and the same measurement window for every candidate. Write acceptance criteria before you look at any output, grade with the same rubric, and record retries and failures rather than only successful runs. Published benchmark scores rarely transfer to your data, so treat them as a screening filter for which models to test.