Return to overview
AI in Recruitment
Remo Vloet7 min.

From Assumptions to Evidence: AI You Can Actually Verify

Every summary line and every extracted field in Simply points back at the moment in the audio it came from. Click it, hear it, and check it yourself.

On this page

“That candidate didn’t sound convinced.” That’s what you tell your colleague after a call. But what if your colleague asks: “Why not convinced? What exactly did she say?” And then you’re standing there. With a feeling. Without proof.

In recruitment, we make decisions every day based on gut feeling. That candidate isn’t motivated enough. That candidate doesn’t fit the culture. That candidate won’t stay long. These are conclusions we draw based on impressions. And sometimes those impressions are right. But sometimes they’re not.

What if you could back up those impressions? Not with a machine’s opinion about the candidate, but with the sentence they actually said and the second of audio it came from.

The difference between a feeling and evidence

A feeling is subjective. “I thought the candidate was somewhat distant.” That might be true. But it could also be that the candidate is introverted, or was nervous, or simply had a bad connection. A feeling is an interpretation coloured by your own experiences, expectations, and even your mood that day.

Evidence is checkable. “She said twice that she would only move for a lead role — at 12:30 and again at 34:10, and here are both.” That’s something you can work with. Something you can share with your colleague. Something you can put in front of a client.

That is what AI summaries plus transparency get you in Simply. The summary is written against the format you chose for that conversation type, and every line of it is anchored to the moment in the audio it came from. The claim and its source travel together, so nobody has to take either on trust.

What the AI actually produces

Worth being precise here, because “insights” is a word vendors use to mean almost anything. Four concrete outputs, and then the list of things that are deliberately not among them.

A summary per conversation type

An intake, a first interview and a client call are different conversations and get different summaries. The format is yours to define, so the summary answers the questions your process actually asks rather than a generic template.

Extracted values as proposals

Notice period, salary expectation, availability, location, the custom field you added last month: these are pulled out of the conversation and arrive as proposed changes on the record, each with a confidence score and the line it came from. You approve or reject. Nothing is written without that.

A transcript that stays attached to the audio

Every sentence of the transcript is anchored to the moment it was spoken. Verifying a number a candidate gave you is a click, not a hunt through forty minutes of recording.

Search across your own conversations

“Who told me they would relocate for the right role” is a question you can ask of the conversations themselves. Retrieval runs behind your permissions, so a record you may not open cannot inform an answer you get.

What is not on that list

No tone analysis. No energy tracking through a conversation. No timing of pauses, no reading of hesitation, no inference about assertiveness, empathy or how somebody holds up under pressure. Those numbers are easy to demo and hard to defend, and the product does not produce them. If a vendor offers you a personality read from a recorded interview, ask them to show you the validation study before you ask them for the price.

Every line points at its source

Here is the part that changes how the output gets used: it is clickable. Not “the candidate seemed unsure about the commute” — which is an opinion a machine has no standing to hold — but the sentence she actually said about the commute, with the audio behind it.

That changes the conversation with your colleagues. Instead of “I thought she seemed unconvinced,” you say: “Listen to minute 23. She says it herself.” That’s a substantiated discussion rather than an opinion contest.

And it changes the conversation with your client. You present a candidate with a summary whose every claim can be played back. When a client challenges one, you answer it in a click instead of promising to check.

What it will not tell you about your own interviewing

Earlier versions of this page promised a mirror: how long you talked versus the candidate, how many of your questions were closed, how often you interrupted. Those numbers are gone, and it is worth saying why rather than quietly dropping them.

Scoring a recruiter from their own recorded calls means reading individual conversations at scale and turning them into a grade. That is a surveillance product wearing a coaching label, it is hard to defend to the person being graded, and a works council is right to ask about it. So there is no talk-ratio number, no question-quality score and no interruption count.

What is left is the thing that actually worked anyway. Open the call with your mentor, replay the passage, and argue about the real thing instead of about a metric. Coaching happens in the record, where opening it requires permission to open it and leaves an entry in the audit log.

Team level: aggregates, and nothing under them

Across a team, the reporting layer is dashboards over your own objects: jobs, applications, placements, and whatever else your data model holds. Counts, rates and trends.

  • Where the pipeline is stuck, and how long things have been sitting there.
  • What converts, per job, per client or per type of work.
  • Where the placements you actually made came from.
  • Open jobs, live applications and outstanding tasks per desk — without opening anybody’s records.

And the constraint that defines the whole thing: there is no drill-through from a chart to the rows behind it. You can see that first-round drop-off rose this quarter. You cannot click the bar to get the eleven candidates behind it, and there is no ranking of your recruiters by how they talk. If you need the records, you open them where the permissions apply in full.

From evidence to better decisions

The value is not in reading a summary. It is in being able to check it before you act on it.

  • The candidate who seemed unsure about the role? Play the passage. Maybe it was not uncertainty but thinking it through, and that is a different signal entirely.
  • The candidate whose answers looked perfect on paper? Read what they actually said, not the tidied version. The summary points at both.
  • Torn between two candidates? Put the two summaries side by side and check the claims that would decide it. The judgement stays yours; what changes is that it rests on sentences rather than recollections.

Privacy and ethics

Recording conversations touches sensitive ground, and two principles do most of the work here. First: output is traceable. The AI does not make claims you cannot check, and where it cannot point at a source it does not make the claim. Second: output is descriptive rather than judgemental. Simply does not say “this candidate is not suitable.” It says what was said, and the interpretation stays with you.

The rest is boring and should be. Ask consent before you record. Agree a retention period and hold to it. The platform and your candidate data are hosted in the Netherlands, on infrastructure Simply runs, and Simply is ISO 27001 certified. AI processing runs in European regions or on your own key with your own provider; the models themselves come from OpenAI and Anthropic, which is said plainly because that is the first thing a security review asks. Your data is not used to train models, and candidates can ask to see what has been recorded about them.

From evidence to action

A summary nobody acts on is a slower set of notes. The connection to the next step is where the time actually comes back.

The proposals are the mechanism. An extracted notice period is not a line in a report — it is a change waiting on the record for you to approve, next to the sentence that produced it. Approve it and the record is current; reject it and nothing happened. That single loop is what keeps a database clean without anybody scheduling a data-quality project.

For a team lead the honest version is narrower than it used to read on this page. The dashboard tells you that applications keep stalling at the same stage, or that one desk is carrying twice the load of the desk beside it. It does not tell you that a particular recruiter is weak at the salary conversation, because it never reads their conversations to find out. It tells you which desk is worth asking about. The asking is still your job, and the answer comes from sitting down together and opening the call.

Evidence as a quality safeguard

Anchoring output to a source is also how you keep the AI honest. When every claim is traceable to a specific moment, you can judge whether it got the context right, and where it did not you reject the proposal rather than correcting a record after the fact. The check happens before the write, not after it.

That is the difference between a tool that tells you to improve your conversations and a tool that shows you what was said, when, and by whom — and then waits for you to decide what it means.

Share this post

About the author

Remo Vloet

Remo Vloet is a co-founder of Simply, the AI Operating System for recruitment agencies: inbox, meetings, sourcing, CV parsing, search and matching, documents and automation in one system. With a background in building complex software, he contributes to the technical vision behind Simply.

LinkedInArticles by Remo Vloet

Frequently asked questions

Can the AI get something wrong?

Yes, and that is why nothing lands on a record without a human approving it. Every summary line and every extracted value carries a link back to the moment in the audio it came from, so checking a claim is a click rather than a re-listen. Treat the output as a draft with its source attached, not as a verdict.

Does this end up in my CRM automatically?

The summary is attached to the conversation, and the values pulled out of it arrive as proposed field changes on the record they belong to. You approve or reject each one, and the approval, the value and the name of the person who approved it go into an append-only audit log in the same transaction.

Does it work on a five-minute call as well?

A summary of a five-minute call is a short summary, because there is less in it. The anchoring works the same either way: whatever the summary does say is traceable to the second of audio it came from, which matters more on a short call than length ever does.

Can I use this for coaching?

Yes, in the conversation rather than in a chart. A mentor and a recruiter can open the same call and replay the passage they disagree about, provided the mentor has permission to open that record. What does not exist is a per-recruiter score to coach against: there is no talk-ratio ranking and no conversation-quality grade, because scoring people from their calls is a different product and not one we built.

Related articles

Chase less,place more.

Find out how Simply can completely evolve your workflow.No slides, just product.