What a 10-minute voice interview can — and can't — tell you
We build voice interviews, so this is the post where we say plainly what the modality is good for — and the part of it we deliberately throw away. Voice reaches a depth a form never will. It also carries a temptation that has skewed hiring for as long as people have been interviewing people: to judge how someone sounds instead of what they said. The honest position is not to pretend the temptation isn’t there. It is to name it, and then build so the machine cannot act on it.
What voice adds
A written application is a monologue the candidate rehearsed. A conversation is not. Three things happen in voice that a form and a one-way video recording cannot do:
Adaptive follow-ups. The next question depends on the last answer. When a candidate says “we cut scope to ship on time,” the interview can ask the question that actually matters — “what did you drop, and what did that cost you?” — instead of moving to the next item on a fixed list. A form asks everyone the same static questions. A conversation follows the thread, which is where the real signal usually lives.
Spontaneous evidence. People rehearse their headline story. They do not rehearse the third follow-up. Ask someone to go one level deeper than they prepared for and you get the unpolished specifics — the actual constraint, the tradeoff they made, the failure case and what they did about it. Those details are hard to fake in real time, which is exactly what makes them worth scoring. A written answer gives you the version the candidate wanted you to see. A conversation gives you the version underneath it.
Lower friction. Talking is easier than writing, for most people and in most languages. There is no essay to draft, no form to grind through, no app to install and no account to create. One browser link, whenever it suits them, about ten minutes. That matters for completion — the candidates who quietly abandon a long written application are often the ones you most wanted to reach. Lower the cost of showing up and more people show up.
What voice tempts you to judge
Here is the uncomfortable half. The same audio that carries a candidate’s answer also carries their accent, their fluency, their pace, and the surface confidence of their delivery. And decades of research say those signals move human interviewers — while saying nothing about whether the person can do the job.
In a controlled study of interview judgments, an applicant’s accent and name shifted how favorably they were rated: an applicant with an ethnic-sounding name speaking with an accent was viewed less positively than the same profile without those cues, and those judgments fed straight into the hiring decision (Purkiss et al., 2006). None of that is competence. It is the sound of a voice doing work that only the content of the answer should be doing.
This is the trap of the modality, and it is worth stating without flinching. The richest part of voice — that you hear the person — is also the part that leaks bias. A tool that listens to a candidate and lets what it hears touch the score has not modernized the old problem. It has automated it, and scaled it to every candidate at once.
The line we draw
So we draw a hard line, and it is architecture, not a promise. Hure transcribes the conversation and scores the words. It never scores the sound.
Concretely, that means no voiceprints, no tone or emotion analysis, no “confidence” detection, no accent scoring — none of it, anywhere in the system. The audio exists to be turned into text, and the text is what gets judged against the rubric. The interviewer does not have a model of how a candidate sounded, because it was never built to. What it has is what they said.
This is not only a design preference; on emotion it is the law. The EU AI Act flatly prohibits AI that infers emotions in the workplace — Article 5(1)(f) puts workplace emotion recognition off the table entirely, not as a box to check but as a line you may not cross. Any interview tool that markets reading a candidate’s tone, stress, or “enthusiasm” is on the wrong side of it. We wrote up what that rule requires, and what NYC asks alongside it, in AI interviews and the law.
The reason we can promise the line held is the same reason it is worth promising: every score cites a moment in the transcript. If a rating pointed at how someone spoke rather than what they said, there would be no citable line behind it — the discipline of anchoring each score to the candidate’s words is what makes it auditable that sound never got a vote.
What 10 minutes is enough for
The other honest thing to say is where the modality stops. Ten minutes is a first round, not a hiring loop, and we would rather undersell it than the reverse.
It is enough time to cover a sealed rubric’s competencies to evidence depth — to ask each one, follow the thread with a real question or two, and come away with a citable answer per competency. It is not enough for a system-design deep-dive, a live coding session, or the kind of long, winding conversation a final round is for. Trying to stretch it into those would produce thin evidence dressed up as a thorough read, which is worse than admitting the boundary.
That boundary is deliberate, because it maps to the one stage where the math actually hurts: the first-round screening call. That is the interview you cannot afford to give every applicant by hand, and the one a consistent, evidence-backed conversation can replace without loss. We ran the numbers on why that stage is the one worth automating in the high-volume screening math. A 10-minute voice interview replaces the screening call. It does not replace the onsite, and it is not meant to.
The close
Voice for depth, transcript for judgment, human for the decision. Use the conversation for what it is uniquely good at, score only the words it produced, and leave the call to a person. The sound is how you reach the substance — it is not the thing you get to judge.