A receding colonnade of identical stone columns in raking light, one column visibly cracked and repaired.

How AI Is Actually Used in Executive Candidate Research

Most writing about AI in recruiting operates at the tool level: which product, which features, which integration. That framing skips the part that determines whether any of it works.

The useful question is narrower. In executive candidate research specifically — assembling a pool, qualifying people against a spec, producing something a client can act on — what does a language model actually do well, what does it do badly, and what has to exist around it before its output is safe to put in front of anyone?

The answer is not flattering to either side of the usual argument. The models are more capable than the skeptics allow, and far less sufficient on their own than the demos suggest.

What the model is genuinely good at

Reading unstructured text at volume. The bulk of candidate research is prose: announcements, filings, biographies, conference programs, company pages. A model reads all of it quickly and consistently, and consistency is not a small thing — a human researcher on hour six of a mapping exercise is measurably not the same reader they were on hour one.

Normalizing things that mean the same thing. “VP Operations,” “Operations Director,” “Head of Manufacturing Operations” describe overlapping jobs across different company sizes. Reconciling those into comparable scope is exactly the kind of fuzzy matching that models handle better than rules do.

Extracting structure from prose. Turning a paragraph of announcement text into fields — role, company, scope, date, source — is reliable, mechanical, and enormously useful, because everything downstream depends on the data being structured.

Generating hypotheses about where to look. Ask which adjacent sectors share a given regulatory constraint set, and you get a genuinely useful list to go searching in. That is a real contribution to the discovery pass, which is mostly a question of knowing where to look.

Drafting the summary once the facts are settled. With verified inputs, writing the entry is the easy part.

Notice what those have in common: each one is a transformation of text that is already in front of the model. That is the zone where this technology is strong.

What it is bad at, and why it matters here

The failure modes are specific, and each one is dangerous in a different way.

It will fill gaps. Asked what someone’s scope was, a model will produce a plausible answer whether or not the evidence supports one. It is not lying. It is completing a pattern. But a fabricated scope claim is indistinguishable, on the page, from a sourced one.

It collapses identities. Two executives with the same name in the same sector will be merged into one person with a remarkable career. This is the most damaging error in candidate research, and it tends to happen more readily with a model than without one: a merged record reads as more coherent than two partial ones, and coherence is what these systems are good at producing.

It states inference in the voice of fact. “Was VP Operations while the company completed an ERP migration” becomes “led the ERP migration.” The reasoning is reasonable. The claim is not confirmed. Prose does not carry the distinction unless something forces it to.

Its confidence does not track its accuracy. The well-sourced claim and the invented one arrive in the same register. There is no tell. This is why “have someone review the output” fails as a control — a reviewer reading fluent, uniformly confident text has nothing to catch on.

It is bad at not answering. Asked a question it lacks evidence for, the default behavior is to produce something. In candidate research, “nothing was found” is frequently the correct and most useful answer, and it is the answer models are least inclined to give.

A reviewer reading fluent, uniformly confident prose has nothing to catch on. Fluency is not a signal of accuracy — it is the absence of one.

What the pipeline around the model has to do

Everything above is manageable. None of it is managed by better prompting — it is managed by structure around the model, and that structure is most of the real work.

It has to be reading, not remembering. A model reasoning over documents that were actually fetched is doing research. A model reasoning over what it absorbed in training is doing recall, and recall is where the invention happens.

A claim and its evidence have to travel together. If sourcing is a step that happens later, it is a step that gets skipped under time pressure — and a citation attached after the fact is a citation that can attach to the wrong thing.

Identity has to be settled before facts get combined. Deciding that two records describe the same person is its own judgment, and it needs to be one the system can get visibly wrong rather than one it makes silently while summarizing.

“I did not find this” has to be a possible answer. If the only shapes available are yes and no, every gap in the evidence gets resolved into one or the other. A third option is what stops a missing source from quietly becoming a negative finding.

The check cannot be run by the thing that did the work. A review sharing context with the generation step inherits its assumptions and confirms its errors. It has to start cold, looking for what the first pass got wrong.

Somebody has to count what the check catches. Whether any of this is working is not a question of how good the output looks. It is a question of what the review is still finding, and whether that number is moving.

Why “just use a chatbot for this” underestimates the problem

The common experiment: paste a role spec into a chat window, ask for candidates, get twelve names in thirty seconds. Some are real and well-suited. Some are real and wrong. At least one does not exist. Nothing carries a source, and nothing distinguishes the three groups.

The instinct is to read this as a model limitation. It looks more like a plumbing one. The model did the part it is good at — pattern completion over text — with no retrieval, no identity resolution, no source binding, and no adversarial check, because none of those were present.

Which is the honest answer to “can I do this myself?” The model is available to everyone and it is genuinely capable. What is missing from the chat window is not intelligence but everything listed in the section above, and it is the same work a research function absorbs by hand when it is done properly.

The line that does not move

There is one thing worth being explicit about, because the tooling makes it easy to blur.

Better research does not extend to deciding who to hire. A system can establish what is true about a candidate and which requirements that satisfies. It has no access to the politics of the last client conversation, the current balance of the executive team, or what the board has quietly ruled out.

Any output that ranks candidates by predicted fit has crossed from research into judgment while presenting itself as research. The correct output is a clear, sourced, honestly-flagged picture of who matches — and who is worth a look despite missing something — handed to the person who has the context to decide.

That division is how Cerna is built. We are not apologizing for it — it is the only version of this a recruiter can defend to a client.

Related reading: how the wider talent-acquisition tooling differs, what AI recruiting bundles together across the funnel, candidate research end to end, what happens when you try it yourself in a chat window, where the three-verdict discipline breaks at the screening step, how to screen executive candidates and how to shortlist candidates for an executive role.

Frequently asked questions

What is AI actually good at in executive candidate research?

Reading unstructured text at volume consistently, normalizing overlapping titles into comparable scope, extracting structured fields from prose, generating hypotheses about where to look next, and drafting a summary once the facts are already verified.

What does AI get wrong in candidate research, and why does it matter?

It fills evidentiary gaps with plausible answers, merges two people who share a name into one false but coherent record, and states inference in the voice of confirmed fact — all in the same confident tone as accurate output, with no visible tell.

What's different between an AI tool built for candidate research and one that just answers a prompt about candidates?

A tool built for research fetches and reads real documents, keeps each claim tied to the source it came from, settles whether two records describe the same person before combining facts, and allows 'not found' as a real answer. A prompt-only tool skips that structure and produces fluent text regardless of whether the evidence exists.

Does using more AI in candidate research make mis-attribution less likely?

No. Mis-attribution tends to get worse with more automation, because a system optimizing for a coherent narrative prefers a merged record over two accurate partial ones, and that preference does not go away with better prompting.

How does Cerna itself use AI in candidate research?

Cerna builds structure around the model rather than trusting its raw output: a claim and its source travel together, identity gets settled before facts are combined, and a separate pass reviews the work cold. Cerna researches and sources candidates against a spec and stops there — it does not rank finalists or decide who to hire.