This interview is with Abhishek Shah, Founder, Testlify.
To start, how do you describe your role at Testlify and your focus in skills-based hiring?
I’m the founder and CEO. Day to day, that means I’m closest to two things: what we build next and conversations with heads of talent who are trying to change how their companies hire but are running into internal resistance. The second shapes the first more than people assume.
Testlify is a skills assessment platform. Companies use us to test whether somebody can actually do the work before deciding how much weight to give their CV.
My focus is narrower than the phrase ‘skills-based hiring,’ because that term has been stretched to mean almost anything. What I care about is sequencing. Most companies that say they hire based on skills have simply added an assessment to a process that is otherwise unchanged: the CV is still read first, and the assessment merely confirms what someone has already decided. That is not skills-based hiring; it is a skills-based garnish.
What key experiences led you to build an assessment platform and shape your perspective on skill evaluation?
We were hiring people for my other business, HNRtech, when I discovered a major gap between what candidates claim they have (in terms of skills and knowledge) and the actual skills they possess.
The problem is that by the time we realize it, it is already too late: you will have already hired them, which leads to twice the cost and time spent.
That’s what made me build Testlify.
Based on your JD-to-assessment work, how do you translate a job description into a calibrated set of skills, weights, and test formats?
I start by discarding most of the job description. Most JDs are a committee wish list, so I keep only what the person will actually do in their first ninety days, and then translate each item into something observable. “Strong communication skills” is not testable. “Write the rejection email to a candidate who got through three rounds” is. If I cannot say what good and bad look like on a specific task, it comes off the list, because neither I nor the hiring manager can score it. Then I cap the list at five to eight skills, since beyond that the assessment gets long, candidates drop out, and the weights stop meaning anything.
For weights, I ask the hiring manager one question: “What does a bad hire in this role get wrong first?” Whatever they say without pausing is the heaviest weight. Format follows from the kind of evidence you need. Tool familiarity can be a short multiple-choice screen because it is cheap and it filters. Judgment under ambiguity needs a deliberately vague brief and a look at what the candidate asks before starting. Craft needs a work sample scored against a rubric by someone who does not know whose work it is. The step almost everyone skips is calibration: run three to five people already doing the job through the test before it goes live. If your strongest performer does not come out near the top, the test is wrong, not the performer.
Building on that, what is your go-to 30-day pilot plan for rolling skills assessments into an existing ATS workflow?
As Abhishek:
One role, not a pilot across the company. Pick a req that hires repeatedly, because you need enough volume in thirty days to learn anything.
-
Week one: pick the role and build the test, then calibrate it against three current employees in that job. Also agree on the one number you will judge the pilot on before you start. Usually that is either time to shortlist or the interview-to-offer ratio, and picking it afterwards is how pilots get argued about instead of decided.
-
Week two: run assessments in parallel with the existing process rather than replacing it. Candidates go through your normal screen, and separately they take the test. Nobody sees the scores yet. This is the only chance you have to find out whether your CV screen and your assessment disagree, and they usually do.
-
Week three: flip the order. The hiring manager now sees the score before the CV, on every candidate. This is the part that gets resisted and it is also the entire point, so do not soften it into a suggestion. Keep the ATS as the system of record and push scores in as a field, so nobody has to work in two places.
-
Week four: compare. How many people did the assessment surface that the CV screen would have dropped, and how did those people do in interviews? Then ask the hiring manager one question: Would you go back? If they hedge, the test measured the wrong thing, and you fix the test rather than the process.
Once live, which outcome metrics best prove predictive validity, and how do you set initial thresholds?
Score-to-performance correlation at six months: it is the only metric that tells you the test predicted the job rather than just sorted the applicants.
What design changes have most improved candidate completion rates without sacrificing rigor?
Making the platform user-friendly and easy to navigate.
Candidates already spend time answering, and we don’t want to take any more of their time understanding how to move around our platform.
In practice, how do you detect and mitigate bias or adverse impact across your question banks and scoring?
We vet all questions in-house with our industry veterans.
When hiring teams are skeptical about automation, what single tactic has consistently earned their buy-in?
We don’t sell automation. We sell you the time you spend on mundane tasks.
Every single real-time decision Testlify makes is backed up by a detailed report, so human judgment is still very much necessary.
Looking 12–24 months ahead, which new assessment or ATS capabilities will most change how recruiters work day to day?
The one that changes the day job is an assessment that adapts while the candidate is taking it. Right now, a test is a fixed set of questions, so half of it is wasted on things you already knew after question four. If the test narrows in on where somebody’s actual ceiling is, you get the same signal in fifteen minutes instead of forty-five, and the drop-off problem mostly goes away.
Second is verification of what happened during the assessment, which matters more each month. It is trivial now for a candidate to have a model beside them. So the interesting signal moves from the answer to the process: what got pasted in, how the work was built, and where they hesitated. That is what we spend a lot of our engineering time on at Testlify, because a score nobody trusts is worse than no score.
The third is less exciting but will change recruiters’ days the most: assessment results sitting natively in the ATS as a sortable field, rather than a PDF attached to a profile. Most recruiters still copy scores between two systems. Remove that, and you remove an hour a day.
What I do not expect is AI making the hire-or-no-hire call. That stays with a person, and the regulation arriving over the next two years will make sure of it.
Thanks for sharing your knowledge and expertise. Is there anything else you'd like to add?
One thing: its the part Id want a reader to leave with. Every company I talk to that says skills-based hiring did not work for them has done the same thing: added an assessment to the front of a process that nobody else changed. The CV is still read first, the manager still interviews the way they always have, and the assessment becomes a step candidates complain about that changes nothing. Then the conclusion is that assessments do not work, when what did not work was bolting one onto an unchanged funnel.
The other thing I would say to any talent leader reading this is to be honest about what you are optimising for. Plenty of hiring processes are built to be defensible rather than accurate, and those are different goals. If a bad hire from an impressive CV is easier to explain to your board than a good hire from an unexpected one, the process will keep choosing the CV, and no tool fixes that.
Im happy to take follow-up questions if anything needs expanding.