What the US AI talent market actually looks like
The US AI talent market is not one market. It is a narrow, severely constrained research and frontier-model tier sitting on top of a much larger, far less constrained applied tier — and most organizations are trying to hire from the first when they need the second.
That mismatch, rather than absolute scarcity, is what makes AI hiring feel impossible. The population of people who can train a frontier model from scratch is genuinely tiny and effectively unhirable for most employers. The population who can take a foundation model, build a reliable production system around it, and keep it running is far larger, far more reachable, and growing quickly. Confusing the two produces a twelve-month search for a role that did not need to be specified that way.
The specification test. Before concluding that AI talent is unobtainable, write down what the person will actually do in their first year. If the honest answer is “integrate models into our product and make it reliable at scale,” you are hiring from the applied tier and the market is much better than the headlines suggest. If the answer is genuinely “advance the state of the art,” the market is as difficult as you have been told, and you should plan accordingly rather than hoping.
Five tiers, five completely different markets
These tiers are routinely collapsed into one requisition. They should not be, because their supply, pay and mobility characteristics have almost nothing in common.
| Tier | What they do | US supply | Mobility |
|---|---|---|---|
| Frontier research | Novel architecture and training-methodology research at the edge of published work. | Very small — low thousands nationally, concentrated in a handful of labs and universities. | Very low. Compensation is rarely the binding constraint; access to compute, data and collaborators usually is. |
| Applied research | Adapting published methods to a specific domain, fine-tuning, evaluation design. | Small but real — tens of thousands, spread far more widely than tier one. | Moderate. Responsive to problem interest and to genuine data access. |
| ML / AI engineering | Building, serving and maintaining models and model-backed systems in production. | Large and growing fast. The deepest pool in the market by a wide margin. | High. This tier behaves like senior software engineering, because that is largely what it is. |
| ML platform & infrastructure | Training and inference infrastructure, GPU orchestration, evaluation and deployment pipelines. | Moderate, and the most persistently under-hired tier relative to need. | Moderate to high. Frequently reachable from adjacent distributed-systems and SRE populations. |
| AI product & safety | Product management for model-backed features, evaluation, red-teaming, governance and risk. | Small but expanding quickly, and drawn from unusually varied backgrounds. | High. The least title-standardized tier in the market, which makes it the hardest to search for and the easiest to map. |
The most common scoping error. Writing a tier-one specification for a tier-three job. A requisition asking for publications at major conferences, distributed training experience at scale, and production ownership of a customer-facing service describes perhaps a few dozen reachable people nationally. Dropping the publication requirement — which the actual work does not need — can expand the reachable population by an order of magnitude without lowering the bar on anything that matters.
Where US AI talent actually sits
Concentration in this market is extreme by US standards. The Bay Area holds a disproportionate share of the frontier and applied research tiers — not because the people were born there, but because the labs, the compute and the funding are. Seattle follows, largely on the strength of cloud and large-scale infrastructure. New York has built a substantial applied and AI-product population on the back of financial services, media and healthcare. Boston is strongly weighted toward life-sciences and academic-adjacent research. Austin, Denver, San Diego, Pittsburgh and the Research Triangle each hold meaningful, specific pockets.
Three things follow from that, and they matter more than the ranking itself.
- Density beats cost in this market. A metro 25% cheaper that holds one fifth of the qualified population is not cheaper, because the search takes three times as long and fails more often. In markets this tight, talent density should dominate the location decision.
- Remote widens the applied tier substantially and the research tier barely at all. Tier-three engineers are widely distributed and comfortable remote. Tier-one researchers cluster around compute, collaborators and institutions, and remote work does not relocate those.
- The secondary metros are specialized, not smaller versions of the primary ones. Pittsburgh is not a small Bay Area; it is a distinct population with distinct strengths. Mapping a secondary metro against a Bay Area comparator set produces a misleading answer in both directions.
Why published rankings are the wrong tool here. Metro rankings count job postings or self-reported profile keywords. Neither measures how many people could actually do your job, and in AI the keyword noise is worse than in any other function — profile self-description has run well ahead of demonstrated capability. Counting the population against a defined capability standard is a mapping exercise, and in this market it is the only method that produces a number you can plan against.
How AI compensation behaves differently
AI compensation is not simply higher than comparable software engineering pay. It is structurally different in four ways, and each one breaks a standard benchmarking assumption.
Equity dominates at the top
In the research tiers, equity frequently exceeds base and bonus combined. A benchmark built on base salary does not merely understate these packages — it ranks them in the wrong order.
The distribution is bimodal
A small set of employers pays far above everyone else for the same nominal title, producing two clusters rather than a bell curve. A median across both describes nobody, and a peer set that mixes them produces an unusable band.
Bands age in months
In tiers one and two, a benchmark more than two quarters old is materially unreliable. Annual survey cycles cannot track this, which is why primary research earns its cost here more clearly than almost anywhere else.
Non-cash terms are decisive
Compute budget, publication freedom, data access and problem selection genuinely move decisions in the research tiers. Employers who can only compete on cash tend to lose candidates they had assumed were priced in.
The practical consequence: benchmark AI roles against a peer set drawn from the tier you are actually hiring from, on total compensation, refreshed at least every six months. The general method is set out in our US compensation benchmarking guide; the difference here is cadence and peer-set discipline, not technique.
The adjacent populations most employers ignore
In a market this tight, the reachable population is usually several times larger than the exact-match population — if you are willing to define the role by capability rather than by prior title.
Distributed systems engineers
The scarcest competence in ML platform work is running large distributed systems reliably, not knowing model internals. That population is large, and the model-specific knowledge is a two-to-three-month ramp.
Quantitative finance
Deep applied statistics, production modeling under real consequences, and a strong evaluation culture. Frequently overlooked because the domain language differs, not the capability.
Computational science and bioinformatics
Large-scale numerical computing, GPU familiarity and rigorous experimental design, often with more methodological discipline than the tier-three average.
Data engineering
Model quality is a data-pipeline problem far more often than a model-architecture problem. Strong data engineers move into ML engineering more successfully than their interview performance usually predicts.
Academic post-docs outside CS
Physics, statistics, computational neuroscience and operations research produce people with the mathematical foundation and the research habits. The gap is engineering practice, which is teachable.
Senior software engineers with evaluation instincts
For AI product and safety work, the binding competence is rigorous evaluation design and judgment about failure modes. That is found in senior engineering and in testing-heavy disciplines, not only in ML backgrounds.
How to use this properly. Adjacency is not a lowering of the bar; it is a correction of the search. The disciplined version is to write down which capabilities are genuinely non-negotiable on day one, which can be acquired within ninety days, and then to count both populations. If the adjacent pool is four times the exact-match pool — which in AI engineering it very often is — the constraint was the specification rather than the market. The same diagnostic is set out in skills that are getting scarce.
Why AI teams leave, and what actually holds them
Attrition in AI teams is high, and compensation is a less reliable explanation than employers assume. Four patterns recur across US organizations.
- The model never shipped. The most common reason strong applied people leave is that their work does not reach production. A year of prototypes that nothing depends on is a credential problem for them, and they know it.
- Compute and data access. In research tiers, the constraint that drives departures is usually infrastructure rather than salary. An employer that cannot supply compute is not competing, however well it pays.
- The vesting cliff. Equity-heavy packages concentrate mobility at predictable moments. Teams where several people joined in the same quarter have a correlated retention risk that nobody has diarized.
- Organizational placement. AI teams buried three layers below the decision they are meant to inform lose senior people fastest. This is a reporting-line problem, and no compensation adjustment fixes it.
All four are observable from outside, which makes them useful in both directions. Competitor teams showing these patterns are reachable, and your own team showing them is a retention problem you can act on before it becomes a search.
How to build an AI hiring plan that survives contact with the market
Assign each role to a tier
Before writing a specification, decide which of the five tiers the role sits in. Most organizations discover that two thirds of what they had called AI roles are tier three, where the market is far more tractable than they believed.
Separate day-one requirements from ninety-day requirements
Write both lists explicitly. The second list is where the adjacency argument lives, and it is the difference between a 40-person and a 400-person addressable market.
Count both populations
Map the exact-match and adjacent pools in your target metros. This is the step that converts an argument about difficulty into a number you can plan against.
Benchmark against the right tier
Use a peer set from the tier you are hiring from, on total compensation, refreshed at least twice a year. Mixing tiers produces a median that loses offers at the top and overpays at the bottom.
Sequence approaches against vesting
In equity-heavy populations, timing is a large part of reachability. Knowing where individuals sit in their vesting schedule turns a cold approach into a well-timed one — the basis of effective pipelining in this market.
Fix what makes the role hold
Reporting line, compute access and a credible path to production do more for both attraction and retention than the last 10% of cash. Establish them before the search, not after the second decline.