When an assistant answers a question about pay, there are two very different things that could be happening underneath. In the first, it is reconstructing a plausible figure from patterns it learned during training. In the second, it is asking a live system for a record and handing back what it received. The answers look identical.
That is the whole subject of this guide. Not whether AI is useful for compensation work, which it plainly can be, but how to tell which of those two things produced the number in front of you, and what changes when it is the second one.
What you will find in this guide
- Two different things an assistant can mean by knowing
- What a model holds about pay
- The cutoff, and why you cannot see it
- What drift costs, measured
- Where model memory is the right tool
- What a live connection changes
- Which part of the answer came from where
- Checking which one you got
- How LaborIQ supports live compensation answers
- Frequently asked questions
Two different things an assistant can mean by knowing
A model does not store facts the way a database stores rows. Training adjusts the weights of a network so that it produces text consistent with what it has read. Where something was stated often and consistently, the reconstruction is close to reliable. Where it was stated rarely, or has changed since, the reconstruction is a plausible guess presented in the same voice.
| CRITICAL DISTINCTION Retrieval returns a record that exists. Generation produces text that fits. A model answering from training is always doing the second, even when the result happens to match a real figure. That is why the same question can produce a slightly different number tomorrow, and why the number cannot be traced. |
The distinction matters most for information that changes. A definition holds still, so a reconstruction of it stays correct. A salary does not hold still, and a reconstruction of one is anchored to whenever the reading stopped.
There is a second consequence that is easy to miss. Because generation is shaped by how a question is phrased, the same assistant will give you a different figure for the same role depending on whether you asked about a “senior analyst” or an “analyst, senior level”. A retrieval system treats those as one lookup or fails to find either. A generative one treats them as two slightly different prompts and produces two slightly different plausible answers, with no indication that anything varied.
What a model holds about pay
It is worth being fair about what is genuinely there. A model trained on a large body of text has absorbed a great deal of real compensation knowledge: how ranges are structured, what percentiles mean, why geography matters, how a pay equity review is run, what questions a compensation analyst asks. That knowledge is durable, because the concepts have not changed.
What it does not hold is the current value of anything. The structural understanding and the numeric content age at completely different rates, and the model presents both with equal assurance. This is why an assistant can give genuinely good advice about how to approach a pricing exercise and a poor figure to start it from, in the same reply.
| What the model holds | How well it ages |
| How compensation structures work | Well. The concepts are stable |
| What the terms mean | Well, with occasional drift in usage |
| How to approach an analysis | Well. Method changes slowly |
| What a role pays | Poorly. It is a value, and values move |
| What a market looks like now | Not at all. There is no now in the weights |
This split explains a pattern people notice and misread. An assistant that has just given an unusually clear explanation of range penetration will, in the next breath, supply a midpoint that is well off. The quality of the first answer is taken as evidence for the second, when the two are drawing on completely different parts of what the model absorbed. Competence with concepts says nothing about currency of values.
The cutoff, and why you cannot see it
Every model has a point beyond which it has read nothing. That boundary is invisible from inside the conversation. The assistant does not experience a gap, because there is no record of what it did not see, and so nothing prompts it to hedge.
This produces a specific and easily missed behavior. Ask about a market condition after the cutoff and the model does not decline. It answers from the last state it knows, in the present tense. Nothing in the phrasing signals that the answer describes a different period from the one you asked about.
| IMPORTANT CAVEAT Telling an assistant the current date does not update what it knows. It changes what is in the conversation, not what is in the model. An assistant that correctly repeats today’s date back to you can still answer your next question from a market that no longer exists. |
The invisibility is worth dwelling on because it is not a bug that will be patched away. A model has no memory of the reading it did, only the residue of it in its weights, so there is no internal list of topics it stopped learning about. Asking an assistant whether its information on a subject is current invites another generated answer, not an inspection. The check has to come from outside.
What drift costs, measured
The gap is not hypothetical, and it can be measured. The chart below indexes the US Employment Cost Index for wages and salaries, and shows what happens to a benchmark that stops being refreshed while the market keeps moving.
| Time since the snapshot | How far behind the aggregate it sits |
| One year | 4.1% |
| Two years | 7.7% |
| Three years | 11.1% |
Two things about those figures deserve care. They describe aggregate wage growth across private industry, so they tell you how far a frozen figure sits below the average, not below the market for your particular role. Individual job families move faster and slower than the aggregate, and the roles that are hardest to hire for are usually the ones moving fastest. So the aggregate is best read as a floor on the drift rather than as an estimate of it.
The second point is about compounding. The gap does not open evenly. It widens with every year the snapshot stays fixed, which is why the difference between a benchmark refreshed annually and one refreshed occasionally grows rather than staying constant.
Where model memory is the right tool
The case against using an assistant for pay figures is not a case against using one at all. The structural knowledge is real and it is immediately useful, provided the numbers come from somewhere else.
| Task | Model memory alone | Why |
| Explaining what compa-ratio measures | Fine | A stable definition |
| Drafting a pay conversation script | Fine | Language, not data |
| Structuring a benchmarking project | Fine | Method is durable |
| Listing what to ask a data vendor | Fine | Reasoning, not values |
| Naming what a role pays in your metro | Not fine | A current value |
| Telling you whether a market moved | Not fine | Requires a now |
The pattern in that table is simple enough to remember. Ask the assistant about how things work and it is a capable colleague. Ask it what something costs today and it is guessing, fluently. We covered the wider risks of relying on generated compensation figures in an earlier piece.
One qualification keeps this table honest. Fine is doing real work in the left column, and it means useful as a draft rather than correct without review. An assistant explaining a concept can still get a detail wrong, and a benchmarking plan it drafts will still need a compensation professional to check it. The distinction the table draws is between tasks where a knowledgeable draft is genuinely valuable and tasks where the output is a number that will be acted on.
What a live connection changes
A connected assistant works differently at the point where a figure is needed. Rather than producing a number from learned patterns, it issues a request to an external system, receives a response, and reports it. The mechanism is covered in our guide to what an MCP server is.
Three things change as a result. The figure is current, because it is fetched at the moment of asking. It is traceable, because a retrieved value can carry its source and the date it was computed. And it is stable, because asking twice returns the same record rather than two independent reconstructions.
| Property | Answering from memory | Answering from a connection |
| Age | Fixed at the training cutoff | As current as the source |
| Traceability | None. There is no record to point at | The source can travel with the figure |
| Stability | Rewording can change the figure | The same request returns the same record |
| Coverage | Broad, thin where material was scarce | Whatever the connected source holds |
| Failure mode | Answers anyway | Can report that it has nothing |
That last row is the underrated one. A system that can say it has no data for a role is more useful than one that will always produce something, because it converts a silent error into a visible gap.
What a connection does not change is worth stating alongside what it does. It has no effect on whether the right job was matched, since that decision is made before the request is issued. It has no effect on the confident register the answer is written in. And it inherits every limitation of the source it is connected to, so a connection to thin or poorly matched data produces well-sourced answers that are no better than the data behind them.
Which part of the answer came from where
In practice a connected answer is rarely purely retrieved. The figure comes from the source; the sentence around it, the interpretation, and any comparison or advice are still generated. This mixture is what makes connected answers easy to over-trust.
| COMMON FAILURES The figure is retrieved and correct. The sentence built around it adds a comparison, a trend claim or a recommendation that came from the model rather than from the source. Because the figure is sound, the whole paragraph inherits its credibility, including the parts that were generated. |
The habit that protects against this is narrow and worth building. Treat the figure and the framing as two separate claims with two different levels of support. The number can be checked against the source. The interpretation around it cannot, and it deserves the same skepticism you would apply to any unsourced reasoning.
A short worked example makes the split visible. An assistant returns a retrieved median for a role, then adds that the figure has risen sharply over the past year and that you should budget accordingly. The median came from the source. The trend claim and the budgeting advice did not, unless the source was asked for a history and it was not. One sentence, two very different levels of evidence, delivered as a single confident paragraph.
Checking which one you got
Three checks separate a retrieved answer from a generated one, and none of them require knowing anything about how the tool is built.
| Check | What to do | What it tells you |
| Repeat it | Ask the same question again, worded differently | A retrieved figure holds. A generated one drifts |
| Date it | Ask when the figure was computed | A source can give a date. Memory gives an adjective |
| Trace it | Ask what it would look like to verify this | A connection describes a record. Memory describes a category |
If all three come back vague, you are looking at model memory, and the figure belongs at the start of a process rather than at the end of one. The wider workflow for turning a market figure into a defensible pay decision is walked through in building a defensible pay decision.
How LaborIQ supports live compensation answers
The gap in the chart above is the gap a refresh cadence closes. LaborIQ recomputes against current market conditions rather than holding a snapshot, so a benchmark is not quietly carrying three years of drift by the time somebody acts on it.
| LABORIQ PLATFORM A Source, Not a Recollection ✓ Recomputed against current market conditions rather than held as a snapshot ✓ The same request returns the same record, so a figure does not move when you reword it ✓ A computation date on the figure, so you can judge it against your own decision ✓ Validated against real company payroll records rather than inferred from text ✓ More than 20,000 unique job titles, including roles a model has read little about → Request a Free Demo at laboriq.co/request-demo |
To check a role against current market data rather than against recollection, start with Salary Answers.
Frequently asked questions
Does the model know what year it is?
Not reliably, and not in the way you would assume. It can often repeat the date you give it, and that does nothing to change what it learned. Telling an assistant the current date updates the conversation, not its knowledge of the market.
Why does it sometimes get a salary roughly right?
Because for common roles in large markets there was a great deal of published material to learn from, and the patterns in it were reasonably consistent. The answer is a good reconstruction of what was widely written. That works until the market moves, and it works least well for the specialist roles where you most need help.
If I paste current data into the conversation, is that the same as a live connection?
For that one conversation, largely yes, and it does not persist. You have supplied the evidence manually, so the assistant is working from it rather than from memory. The next conversation starts without it, and so does everyone else in your team.
Is web search enough to solve this?
It helps with age and it does not solve provenance. A search result can be a blog post quoting a figure from an unnamed source, which is a citation without a chain behind it. What you want is a connection to a system that holds records, not to a page that mentions a number.
How stale is too stale for a salary benchmark?
That depends on the role and the market rather than on a fixed number of months. A useful test is to ask what decision you would make differently if the figure were a few percent higher. If the answer is a real difference, then a benchmark old enough to have drifted by that much is too old for that decision.
Does this apply to every kind of AI tool?
It applies to anything answering from a trained model without a connection to a data source. Tools that retrieve before they answer behave differently, and it is worth knowing which kind you are using rather than assuming from the interface, since both look the same from the outside.
Should we tell our teams to stop asking assistants about pay?
A ban is difficult to enforce and tends to push the behavior out of sight. A clearer rule is that a figure from an assistant is a starting point and never the number that goes into an offer, and that anyone quoting a figure should be able to say where it came from.
