A weathered stone sundial with a bronze gnomon casting a sharp, precise shadow, half its carved numerals worn past legibility.

ChatGPT for Recruiting: What It Can Do for Executive Search

Say you’re filling a VP Operations role at a mid-size manufacturer, and instead of opening a database you open ChatGPT. You paste in the brief: the must-haves, the sector, the rough seniority. You ask for candidates. This is the actual test worth running, because the answer to “can I do this myself” depends entirely on what happens next, not on whether the first response looks good.

The first response looks good

You get a list. Names, current titles, a paragraph on each explaining why they fit. It reads like research. A few names you recognize, which feels like a good sign. You start reading the paragraphs more closely, and this is where the useful part of the exercise actually begins, because up to now you’ve learned nothing except that the model can produce a plausible-looking list, which was never in doubt.

Pick one candidate and ask a follow-up: where does that P&L figure in their bio come from, or what exactly did they run at their last company. Two things tend to happen. Sometimes you get something you can act on: the model correctly pulls a distinction from the text it has access to, the kind of “VP Operations” versus “Director of Manufacturing” scope difference that’s genuinely fiddly to reason about and that a model is often faster at than you’d expect. Other times the answer is confident, specific, and simply asserted; a P&L number that isn’t in any source, a scope that’s been rounded up because it makes a cleaner story. Nothing in how either answer is written tells you which one you’re getting. Both come back as complete, well-formed sentences with the same tone.

Where the list starts to fall apart

Keep pulling on individual candidates and two more problems show up. One: if there are two people with a similar name in the same industry, the model will sometimes hand you a single biography that’s actually a merge of both of them, an early role from one person grafted onto a recent title from the other. Without independently knowing both careers, there’s nothing in the text itself that marks the seam where one biography stops and the other starts. Two: ask a follow-up question and watch the claim quietly grow across the conversation. A title the source confirms as of a specific year drifts, a few questions later, into having started well before that, the tenure stretching backward each time you ask for more detail until what you’re reading no longer matches what the original page said.

None of this is the model being careless in a way you could train out of it by asking more carefully. Producing a confident, complete answer to whatever’s asked is what it’s built to do. There’s no visual difference between a sentence built on something it actually has in front of it and a sentence built on a good guess, and rereading the paragraph a third time doesn’t change that. What changes it is going and checking the specific claim against something you can point to.

What you’d have to add to make it usable

None of this means the exercise was wasted. It means the free version and the version you could defend to a client are different amounts of work, and the gap is specific enough to name.

You’d need to note down where each fact came from as you go, not reconstruct it afterward by asking the model to cite itself. Sourcing added after the fact tends to attach to whichever claim it’s asked about, correct or not, rather than to the claim that actually needs it.

You’d need to check every repeated or common name by hand rather than trust that one biography means one person. The model won’t flag an identity it has silently merged; you’re the only check on that, and it means reading closely enough to catch a career gap or a location that doesn’t line up.

You’d need to be willing to write “couldn’t confirm this” and leave it there, instead of asking again until you get an answer. A model that always answers makes it impossible to tell a verified fact from a filled gap by looking at the output alone.

And you’d need someone who didn’t run the search to look at the finished list before it goes anywhere, specifically hunting for what doesn’t hold up. Checking your own work against the assumptions you used to build it misses things a second, uninvolved reader catches on the first pass.

Do all of that, on every name on the list, and you’ve built something closer to real candidate research, with a language model doing part of the reading. Skip it under deadline pressure and what you have is a plausible-sounding list with no way to know, later, which lines were checked and which were guessed.

What it still won’t tell you

Even a fully checked list doesn’t answer who to hire. Confirming that someone really carried the scope your brief requires is a different question from whether they’re right for this team, this client relationship, this moment, and that second question runs on context no model has: how the last VP Operations hire at this company actually worked out, what the plant floor is missing that nobody currently on the team can cover, what this client will and won’t sign off on. That judgment stays with the person running the search, whether or not a model helped with the reading.

A chat window still turns a real brief into a real, readable answer faster than almost anything else you could open first. What it doesn’t do on its own is keep every line honest afterward: tracing each claim back to something you can open, catching the merge before it goes anywhere, leaving a genuine gap marked as a gap instead of guessed at, putting the finished list in front of someone who wasn’t in the room. Building that in yourself, candidate by candidate, is the part of the research bottleneck a fast first answer doesn’t remove. It just moves the work to after the response instead of before it.

Related reading: what the wider AI recruiting category actually covers, what a language model does well and badly across the research pipeline and what a good executive search shortlist actually contains.

Frequently asked questions

Can ChatGPT do executive candidate research on its own?

It will produce a plausible-looking list of names quickly, but nothing on the page distinguishes a claim pulled correctly from the source text from one that is confidently invented. The test is pulling on one specific claim and checking where it actually came from.

What happens if you ask ChatGPT a follow-up question about a candidate it listed?

Two things happen unpredictably: sometimes it correctly pulls a real distinction from text it has access to, and sometimes a claim quietly grows across the conversation, with a title's start date stretching earlier each time you ask for more detail.

Does ChatGPT flag when two candidates share a similar name?

No. It can silently merge two people into a single biography, grafting an early role from one person onto a recent title from the other, with nothing in the text marking where one person's career ends and the other's begins.

What would I have to add myself to make a ChatGPT candidate list usable?

Note where each fact came from as you go rather than asking the model to cite itself afterward, check every repeated name by hand, write 'couldn't confirm this' instead of asking again until you get an answer, and have someone who didn't run the search review it.

Does a fully fact-checked ChatGPT candidate list tell you who to hire?

No. Confirming that someone carried the scope your brief requires is a different question from whether they are right for this team and this moment, and that judgment runs on context no model has access to.