What compensation benchmarking is in the US context
Compensation benchmarking is the process of establishing what a defined population of people is actually paid, in a defined market, for work of comparable scope — and then positioning your own pay against that evidence.
In the United States, the discipline is harder than in most markets for one structural reason: base salary is frequently a minority of the package. A Senior Staff Engineer in the Bay Area may earn $210,000 in base and $340,000 in total compensation. Benchmarking that person against a base-salary survey produces a number that is not merely imprecise, it is answering a different question. US benchmarking has to be done on total compensation or it is not benchmarking at all.
There is a second structural problem. US job titles are not standardized. “Director” at a 300-person company and “Director” at a Fortune 100 describe roles two organizational layers apart with a compensation gap that can exceed 100%. Any benchmark built by matching titles rather than scope will produce a confidently wrong answer — and confidently wrong is the expensive kind.
The operating test. A benchmark is only sound if you can name the companies in the peer set, state how each comparator role was matched to yours, and show that the compensation figures were verified rather than self-reported. If any of those three is missing, what you have is a reference point, not a benchmark.
The six components of US total compensation
Every one of these has to be priced separately and then reassembled, because candidates evaluate them separately.
Base salary
The only component most surveys capture well, and in senior US roles frequently the least informative. Still the anchor for benefit calculations, severance and internal equity.
Typical share: 45–85% depending on level and sector.
Target and actual bonus
Target bonus is policy; actual payout is reality. The gap between them over three years is one of the most revealing things about an employer, and almost never appears in a published survey.
Typical share: 10–40%.
Equity grant value
RSUs at public companies, options or RSUs at private ones. At public companies the annual refresh grant matters more than the sign-on, because the refresh is what makes the package durable.
Typical share: 0–60%.
Vesting position
Not a value but a timing fact, and the one that determines whether someone is actually reachable. A candidate eleven months from a cliff is a different prospect from the same candidate two months after one.
Effect: determines mobility, not cost.
Benefits and retirement
Employer healthcare contribution, 401(k) match and vesting schedule, and in some sectors deferred compensation plans. In the US the healthcare contribution alone can be a five-figure annual difference between two otherwise similar offers.
Typical value: $8,000–$30,000+ annually.
Location and work model
Geographic pay differentials, remote-work pay policy, and whether the employer localizes pay when someone relocates. Post-2020 this became a negotiating point rather than an administrative rule.
Effect: 0–35% swing on the same role.
Why equity is where US benchmarks break
Equity is the component most often mishandled, and it fails in both directions. Valuing a private-company option grant at its paper strike-price arithmetic overstates it, sometimes wildly. Excluding equity because it is hard to value understates public-company packages by a third or more. Neither produces a usable comparison.
The defensible approach is to value public-company RSUs at the annualized grant value using a trailing average share price, treat late-stage private equity at the most recent preferred round with an explicit illiquidity discount, and record early-stage equity as a percentage of the company rather than a dollar figure. Then state the method on the page next to the number. A benchmark whose equity method is not written down cannot be audited, and an unauditable benchmark will not survive its first challenge from a hiring manager.
Where US compensation data comes from, and how far to trust each source
| Source | Strength | Weakness | Use it for |
|---|---|---|---|
| Federal wage statistics | Authoritative, national, free, methodologically consistent over time. | Occupational categories are broad, data lags by a year or more, and there is no equity or bonus detail at all. | Macro context and cost-of-labor comparisons between metros. Never for setting an individual offer. |
| Posted pay ranges | Current, employer-stated, and now legally required in a growing number of states. | Ranges are often deliberately wide, and the posted band is the policy range rather than where offers actually land. | Establishing the ceiling a competitor is willing to publish, and tracking movement in that ceiling over time. |
| Purchased salary surveys | Structured, leveled, statistically presented, defensible to a compensation committee. | Participant sets are self-selected, submissions are self-reported, and data is typically 6–18 months stale by publication. | Internal band architecture and grade design. Weak for hot or fast-moving roles. |
| Crowdsourced platforms | Fast, current, and unusually good on equity detail at large technology employers. | Unverified, self-selected, and skewed toward a narrow set of well-paid roles at well-known companies. | A directional sanity check. Never as a primary source and never on its own. |
| Primary research | Verified, current, role-specific, and matched on actual scope rather than job title. | Costs money and takes time. Coverage is only as wide as the population you commission. | The roles where the decision is expensive and being wrong is worse than being slow. |
The practical combination. Use federal data for macro geography, purchased surveys for band architecture, posted ranges for competitor ceilings, and primary research for the twenty or thirty roles where the money and the risk actually sit. No US organization needs verified primary data on every role. Most need it on far more than none.
How to run a US compensation benchmark properly
Eight steps. Skipping step three is the most common cause of a benchmark that fails under challenge.
Define the decision first
Repricing an existing population, constructing a band for a new role, and preparing a single counter-offer are three different exercises with three different peer sets. Establish which one you are doing before collecting anything.
Build the peer set deliberately
Name the comparator organizations and write down why each belongs: competing for the same people, similar scale, similar operating complexity, same metro. “Companies in our industry” is not a peer set. A good one usually holds 15–40 named organizations.
Match on scope, never on title
Compare reporting line, team size, budget owned, P&L responsibility and decision rights. A Director running 60 people and $40m is not comparable to a Director running four people, whatever the two business cards say. This step is what separates a benchmark from a title survey.
Collect all six components
Base, target bonus, actual bonus history, equity grant value, vesting position and benefits. A partial collection produces a partial answer that will be presented as a complete one.
Verify against independent sources
Every compensation figure should be triangulated. Where it cannot be verified, mark it as an estimate and show the confidence level rather than quietly blending it into the median.
Normalize for geography and date
Adjust for metro differential and age the data forward to a common reference date. A figure collected fourteen months ago and used raw is not a current benchmark.
Report distribution, not just a median
Give the 25th, 50th and 75th percentile and the sample size behind each. A median drawn from six data points should not be presented with the same confidence as one drawn from sixty, and a single number hides the spread that the actual negotiation will happen inside.
Write down the method
Peer set, matching logic, equity valuation approach, sample sizes, collection dates and known gaps. This is what makes the benchmark defensible when a hiring manager disputes it, and reusable when someone repeats the exercise next year.
Metro differentials and why national averages mislead
The United States is not one labor market. It is several dozen metropolitan labor markets with distinct competitor sets, distinct pay levels and distinct mobility patterns. For the same role at the same scope, the spread between the most and least expensive major US metros commonly exceeds 40%, and in software engineering it can approach 60%.
A national average sits in the middle of that distribution and describes almost nobody. Used to set a band in a high-cost metro it will lose every competitive offer; used in a lower-cost metro it will overpay on every hire and quietly compress the internal structure around those hires.
Three differentials that behave differently
- Cost of labor. What comparable employers in that metro actually pay. This is the differential that matters for competitive offers, and it is a function of local demand density, not local rents.
- Cost of living. What it costs an employee to live there. Relevant to candidate perception and relocation conversations, but a poor basis for setting pay — the two diverge substantially in several US metros.
- Talent density. How many qualified people are actually present. Often the most decisive factor and the one most often omitted. A metro that is 15% cheaper but holds one third of the qualified population is not cheaper once time-to-hire and search risk are priced in.
Where this connects to mapping. Talent density is not available from any compensation survey, because surveys count salaries, not people. It comes from talent mapping — counting the qualified population in each metro. Benchmarking answers what the market pays. Mapping answers whether the market has anyone to pay. Location decisions need both, and organizations that buy only the first routinely site teams in metros where the population does not exist.
What pay transparency has changed for benchmarking
A growing number of US states now require employers to publish a good-faith pay range in job postings. That has changed benchmarking in three concrete ways, none of which were the stated intent of the legislation.
- Competitor bands are now observable. You can read what rivals are willing to publish, track how those ceilings move over time, and see which employers quietly widened their bands rather than raising them.
- Internal equity is now externally visible. When a posted range for a new hire sits above what an existing team member earns, that employee can see it. Benchmarking that ignores the internal population now produces a retention problem, not just a hiring one.
- Multi-state employers are converging. Publishing different ranges for the same remote role in different states is defensible in principle and awkward in practice. Many employers have responded by narrowing geographic differentials rather than defending them individually.
The obligations vary by state, by employer size and by whether a role is remote-eligible. We cover the state-by-state position in detail on our guide to US pay transparency laws.
Seven ways US compensation benchmarks go wrong
- Matching on title. The most common and most expensive error. US titles are not comparable across companies, and a title-matched benchmark reliably produces a plausible number for the wrong job.
- Benchmarking base only. In sectors where equity and bonus are 40–60% of the package, a base-only benchmark is not conservative, it is wrong. It will lose offers while appearing competitive on paper.
- Using stale data on a fast-moving role. An eighteen-month-old survey figure for a role in an actively contested skill area is not a benchmark, it is a historical note.
- Peer sets chosen for aspiration. Benchmarking against companies you admire rather than companies you compete with for people produces bands you cannot fund and did not need.
- Ignoring the internal population. An external benchmark applied only to new hires creates compression, and compression shows up as resignations from people you were not planning to replace.
- Reporting a median with no sample size. A median from five unverified data points and one from sixty verified profiles look identical in a slide. They should never be presented the same way.
- No documented method. A benchmark whose peer set, matching logic and equity treatment are not written down cannot be defended, cannot be repeated, and will be overturned by the first senior stakeholder who disagrees with it.
What a benchmarking deliverable should contain
- Named peer set with the inclusion rationale for each organization.
- Role-matching record showing how each comparator was matched to your role on scope, and where the match is approximate.
- Distribution by component — 25th, 50th and 75th percentile for base, target bonus, actual bonus and equity, each with its sample size.
- Metro breakdown where the population spans more than one US market.
- Total-compensation view reassembled from the components, with the equity valuation method stated explicitly.
- Your current position plotted against the distribution, for both new hires and the existing population.
- Confidence marking on every figure — verified, triangulated or estimated.
- The raw dataset in Excel or CSV, owned by you, reusable without restriction.
Audentia delivers all of the above as a fixed-fee project. Compensation work is usually commissioned either as a layer on a US talent mapping project, where the population is already identified, or as a standalone benchmark against a named peer set. Pricing follows the same project-fee model set out in our US cost guide.