AI Agents Are Now Interviewing Candidates. Should Yours?
Fika Jobs raised $4M for AI agents that interview candidates. The evidence says automate your funnel, not your judgment. Here's the line, and why it matters.
Ernest Bursa
AI should handle the high-volume, objective parts of hiring, such as parsing applications, scheduling, and transcribing and searching interviews, but not the final evaluative decision. Research shows AI succeeds on clear, verifiable tasks 76 to 82% of the time, yet drops to as low as 33% on subjective synthesis, and it reproduces racial and gender bias at scale. The rule for founders is simple: automate the funnel, not the judgment. Keep a human on the call about whether this is the right person.
Should AI interview your candidates? The short answer
In June 2026, a Swedish startup called Fika Jobs raised $4 million to build a platform where an AI agent, not a recruiter, conducts the interview. It is a real product with real money behind it, and it forces a real question for anyone building a hiring funnel this year: how much of your own process do you hand to a machine?
The honest answer is “a lot, but not the judgment.” The interview exists to answer one question: is this specific person right for this specific role? That is the exact task the current generation of AI is measurably worst at. Meanwhile, AI is excellent at everything surrounding the interview: the parsing, the scheduling, the transcribing, the searching, the summarizing. So the line isn’t AI versus no AI. It’s which stage. Put AI where it’s strong and keep a person where the stakes and the weaknesses collide.
This matters because the flood of AI-written applications has made volume unbearable, and automating the interview looks like the only way to keep up. It’s a trap worth naming, because getting the line wrong doesn’t just cost you a hire. It can cost you a discrimination claim and your employer brand.
The $4M bet: what Fika Jobs is actually building
Fika Jobs raised a $4 million pre-seed round led by Luminar Ventures, with participation from Alliance VC and angel money from King (Candy Crush) co-founders Sebastian Knutsson and Riccardo Zacconi, according to TechCrunch’s June 23, 2026 coverage. The founders, brothers Jakob and Alexander Dubois, previously built the social app Gaff.
The product is a video-first hiring platform. A candidate connects their LinkedIn profile, an AI agent powered by Google’s Gemini models generates personalized questions, and the agent conducts a roughly 10-minute video interview. The platform then auto-converts the answers into short clips organized into a live, browsable candidate profile. Instead of applying role by role, candidates maintain a standing profile that employers browse. It’s free for job seekers, and Fika takes 10% of the hire’s first-year salary, which the founders position as cheaper than the 20 to 30% traditional recruiters charge. Over 100 companies are on the waitlist, 50-plus have tested it, and a broader launch is expected in fall 2026, starting in Sweden.
The efficiency case is obvious, and the convenience is genuinely candidate-friendly. But here’s the tell: TechCrunch’s own funding announcement flags that video profiles “introduce real bias risks,” because employers see race, age, gender, appearance, and accent before they evaluate qualifications. When the bias caveat is baked into the launch coverage itself, that’s the whole debate in miniature.
Where AI genuinely wins in the hiring funnel
AI is strong at high-volume, objectively checkable, organizational work. This is the 76 to 82% zone, and you should automate it aggressively.
The same research that indicts AI on judgment (more on that below) found the leading models pass clear, verifiable tasks 76 to 82% of the time. That maps cleanly onto the parts of hiring nobody enjoys doing by hand:
- Triage. Parse and de-duplicate applications, and rank them against objective must-haves like a required certification, work authorization, or years in a specific stack.
- Scheduling. The pure-friction step that burns weeks of back-and-forth. Slow scheduling is itself a real cost, since scheduling delays drive candidate drop-off before you ever meet your best people.
- Reading. Transcribe interviews, summarize a candidate’s history, and search across hours of video and many candidates in seconds.
- Summarizing. Turn a sprawling application and transcript into a tight brief a busy hiring manager can actually read.
None of these are the decision. They are the preparation for the decision. Done well, they give your team hours back and let a human spend their scarce attention on the part that needs a human.
Where AI breaks: the 33% synthesis problem
AI is measurably bad at the exact judgment an interview requires: weighing incomplete, subjective signals into one clear conclusion.
PYX Labs, a research lab sponsored by employee-experience firm Perceptyx, tested seven leading AI models from OpenAI, Google, Anthropic, and xAI across 84 employee-listening tasks, scored against criteria built by psychologists and organizational-behavior specialists, as reported by HR Dive. On clear, verifiable tasks the models passed 76 to 82% of the time. But when they had to interpret open-ended, subjective feedback and resolve it into a coherent takeaway, success dropped to as low as 33%. Synthesis was the single lowest-scoring capability across every model, ranging from 14 to 57%.
Read that failure mode again: “weighing incomplete, emotional, or context-dependent signals and resolving them into one clear takeaway.” That is a working definition of interviewing a human being. The stat that matters isn’t that AI is bad at everything, because it isn’t. It’s that AI is worst at precisely the one thing you’d be asking it to do if you let it run the interview.
For context on how quickly this is spreading, 37% of respondents to a 2025 WTW survey already use AI in performance management, per HR Dive. Adoption is racing ahead of the capability ceiling. That gap is where the risk lives.
The bias tax: what AI screening does at scale
When AI makes the evaluative call, it doesn’t just get judgment wrong. It reproduces discrimination systematically, in a way that’s invisible until someone measures it.
A University of Washington study by Kyra Wilson and Aylin Caliskan tested three open-source LLMs against more than 550 real resumes and over 500 real job listings across nine occupations, running more than three million comparisons using 120 first names. The results, presented at the AAAI/ACM Conference on AI, Ethics, and Society in October 2024:
| Comparison | AI favored | vs. |
|---|---|---|
| White- vs. Black-associated names | 85% | 9% |
| Male- vs. female-associated names | 52% | 11% |
| Black-male vs. white-male names | never | — |
That last row is the one that should stop you. The systems never once favored Black-male-associated names over white-male-associated names, an intersectional harm that isn’t just “race plus gender” stacked together. And this is the same class of finding as the largest deployed-AI-hiring study to date, which found clear racial disparities across millions of real applications. This is not a lab curiosity. It’s what these systems do when you point them at real candidates and let them decide.
The point is not “never use AI.” The same models that score 33% on synthesis score 76 to 82% on verifiable tasks. The point is that the decision-maker is the one seat where these failures become discrimination, and that seat should not be automated.
The trust and legal line
Two risks converge on the moment an AI makes the call: candidates stop trusting you, and regulators start noticing.
On trust, be honest that the convenience is real. Async AI interviews remove scheduling friction, let people interview at 11pm, and can feel less biased than a distracted human to some applicants. But many candidates report feeling dehumanized talking to a bot, uncertain how they’re being scored, and unable to appeal a black-box rejection. The convenience win and the trust risk live inside the same product. The only resolution is transparency about what the AI does and doesn’t decide, which a human-in-the-loop process can promise and a fully autonomous one cannot.
On law, the exposure is growing and jurisdiction-specific. New York City’s Local Law 144 requires an independent bias audit and candidate notice before you use an automated employment decision tool. Illinois’ Artificial Intelligence Video Interview Act requires disclosure, consent, and explanation when AI analyzes video interviews. Colorado’s AI Act adds duties around high-risk AI in employment decisions, phasing in around 2026 and 2027. And the EEOC has signaled that using a biased algorithm doesn’t shield an employer from Title VII liability.
The practical takeaway for a founder: “the vendor’s model did it” is not a defense. If an AI tool rejects protected-class candidates at a disparate rate, you own the outcome. A human-reviewable decision trail isn’t just good ethics. It’s the exact audit artifact regulators and plaintiffs ask for.
What human-in-the-loop actually looks like
Human-in-the-loop means the evaluative decision is an explicit, attributed action by a named person, and the AI’s role is bounded to assistance. It should be a constraint in your system, not a line in your careers-page copy.
This is exactly how Kit’s hiring vertical is built. The funnel is made of typed stages: screening (which can auto-screen on objective must-haves), team_review (with configurable voting, including a score threshold, a require-all-reviewers option, and a veto_auto_rejects flag so a person can overturn any automated rejection), live_interview (scheduled, with a required number of human interviewers), and code_assignment. Every advance or reject is a Decision record whose decided_by must be a human account user. There is no code path for “the AI rejected this candidate.” The system literally cannot record a verdict that isn’t attributed to a person.
AI shows up as tools, not as the decider. Kit’s MCP-based tools, like candidate and application summaries and semantic search across interview transcripts, help your team read faster across many candidates and hours of video. They surface information; they don’t cast the vote. (For how that tooling is built, see MCP for hiring.) That’s the difference between “AI interviewed and rejected you” and “a person decided, with AI helping them read faster.”
And because the candidate portal exposes real, human-reviewable stage progress instead of an opaque score, it’s both the trust answer for candidates and the audit answer for regulators. Note that this is a different problem from candidates gaming AI screens, which we covered in AI interview cheating and redesigning assessments. This article is about who holds the pen, not who’s trying to beat the test.
A stage-by-stage rule for founders
So, should your AI interview candidates? No. But it should do almost everything around the interview. Here’s the line, stage by stage:
| Stage | Let AI do it? | Why |
|---|---|---|
| Parse and de-dupe applications | Yes | High-volume, objective, verifiable |
| Rank against must-haves | Yes, with a human-vetoable auto-reject, never a silent one | Objective checks are AI’s strong zone; the veto catches the bias |
| Scheduling | Yes | Pure friction, no judgment |
| Transcribe and search interviews | Yes | Reading fast, not deciding |
| Summarize candidate history | Yes | Surfaces context for a human |
| The interview evaluation | No | This is the 33%-synthesis zone |
| Every advance or reject | No | Judgment plus legal exposure; keep it human and logged |
That split captures most of Fika’s efficiency without inheriting the 33%-synthesis and 85%-name-bias liabilities of an autonomous interviewer. Add candidate-facing transparency about which is which, so nobody is guessing whether a bot or a person made the call.
The founders reaching for AI interviewers aren’t wrong about the pain. Volume is crushing, recruiter fees are steep, and scheduling is a slog. They’re just aiming the automation at the wrong seat. Point it at the funnel, keep a human on the judgment, and you get the speed without the trust gap. Kit is built on that exact line: the efficiency of AI triage with the defensibility and humanity of human decisions, proven in the data model rather than promised in a pitch. If you want a pipeline that works this way out of the box, explore the role templates or start free.
Related articles
Ready to hire smarter?
Start free. No credit card required. Set up your first hiring pipeline in minutes.
Start hiring free