Interview with Abhishek Shah, Founder, Testlify

Connectively

Connectively connects subject-matter experts with top publishers to increase their exposure and create Q & A content.

7 min read

Interview with Abhishek Shah, Founder, Testlify

© Image Provided by Connectively

This interview is with Abhishek Shah, Founder, Testlify.

For Featured readers who may not know you, how do you describe your current focus and expertise in skills assessment and talent identification?

I run Testlify, which builds skills assessments companies use instead of resumes. My focus over the last couple of years has been narrower than it sounds: getting assessment design right for roles where a good-looking resume and actual competence have drifted apart—especially engineering and support roles, where an AI-polished application makes real screening harder instead of easier.

Most of what I spend my time on isn’t about strategy decks or hiring trends. It’s the assessment itself: the actual task a candidate sits down and does.

If a multiple-choice test can be answered by a chatbot in ten seconds, the score is useless. Rebuilding tests around that reality has become most of my job. Talent identification, the way I use the term, means finding whoever can do the specific job in front of them, not whoever interviews best or clears a keyword filter first.

What were the pivotal experiences that shaped your path to leading Testlify and your philosophy on evidence-based hiring?

I was already running a business before Testlify; that’s really where this one came from. I kept hitting the same wall as a hiring manager myself: the assessment tools on the market were either bloated, fifty-question marathons or generic personality quizzes that had nothing to do with the actual job. Half of them felt built to satisfy an HR checklist, not to tell you whether someone could do the work.

The resentment on both sides was the real signal, more than any single bad hire. Candidates hate a test that eats an hour and tells them nothing about the role. Recruiters hate a score they can’t explain to a hiring manager two rounds later. Once that gap was obvious, it stopped being an observation and turned into something I had to go build.

The philosophy that came out of it is easy to say and harder to run: move the call from a gut read to something you can defend with evidence, not a personality score but something closer to watching someone handle a version of the actual job, a situational read instead of a static one. Candidate experience turned out to matter almost as much as accuracy, because a test people resent gets rushed, faked, or abandoned halfway through, and any of those outcomes tells you nothing real about the person on the other end.

What is the most counterintuitive lesson you’ve learned about using skills assessments to predict on-the-job performance?

The one that surprised me most was that making the assessment more realistic didn’t automatically make it more predictive. We kept assuming that the closer a task looked to the real job, the better it would forecast performance, so we built long, layered simulations of actual work. Complexity went up and precision dropped instead of improving.

What we found instead was that a tightly scoped, thirty-minute task with one genuinely ambiguous decision point told us more than a ninety-minute simulation ever did. The long version mostly measured who had the patience to sit through a long exercise. It rewarded stamina, not skill.

The part that still surprises people who haven’t run into it themselves is that the second-highest scorer often outperforms the highest scorer once they’re actually on the job. The top score sometimes belongs to whoever read the test well, not the role. Once we started asking candidates to walk through their reasoning instead of trusting the score alone, the gap between assessment results and real performance closed more than adding extra questions ever managed.

You built a JD-to-assessment engine at Testlify—what single design decision most improved signal quality, and how did you know it worked?

The decision that mattered most was to refuse to let the AI simply read a raw job description and improvise a test from it. We constrained the input to a handful of things that actually predict what’s needed: skills, experience level, question type, and difficulty. A JD is full of noise: boilerplate, legal language, and aspirational ‘nice-to-haves’ that nobody actually checks for. If you feed all of that in, the AI generates a test that sounds thorough but measures nothing specific.

I knew it was working when recruiters stopped rewriting the tests before sending them out. That sounds small, but it wasn’t happening before. When a generated test needs a rewrite, it means the signal was off enough that a human didn’t trust it. Once that editing step started disappearing on its own, without us telling anyone to skip it, that was the real tell that the narrower input was doing its job.

For teams hiring on a near-zero budget, what is your step-by-step playbook for identifying top talent using free or low-cost tools?

Skip the resume screen if the budget is actually zero — that’s where most cash-strapped teams waste the little time they have.

  1. Post the role somewhere free. A niche Slack community or subreddit usually works better than a job board, and ask for one thing in the application: a short answer to a real problem your team is dealing with right now. You’ll get fewer applicants, but almost all of them will have actually read the post.

  2. Build one work-sample task in a Google Doc or a public GitHub repo: something under an hour that mirrors real work instead of a brain-teaser.

  3. Write down three or four criteria you’re grading for before you open a single submission, so your own bias doesn’t creep in halfway through.

  4. Do a free 20-minute call with anyone who clears that bar and have them walk you through their answer out loud. What people say about their own work usually tells you more than the work itself, and none of it costs anything beyond your time.

When a hiring manager is unsure which skills to prioritize, what rapid 30-minute role analysis do you run to define the core competencies and their weights?

Thirty minutes is enough once you stop asking about traits and start asking about last week. I sit the hiring manager down and ask them to walk through the last real problem someone in this role actually solved, in detail—not what the job description claims the role covers. Ten minutes in, you usually land on four or five concrete tasks instead of a list of adjectives like “strategic” or “detail-oriented.” A task is something you can actually build a test around; an adjective isn’t.

Then I ask one blunt question about each task: What happens if someone is bad at this specific thing? Some answers are “we lose a client within a week.” Others are “nobody notices until the quarterly numbers come in, and by then we’ve course-corrected.” That gap is your weighting, not a guess about which skill sounds most senior on paper. Whatever ties to fast, visible failure gets the heaviest weight in the assessment. Everything else still gets tested but doesn’t dominate the score, and that twenty-minute conversation settles arguments that a job description alone never resolves.

How do you calibrate assessment difficulty and cut scores to reduce false negatives without lowering the hiring bar?

Most teams calibrate difficulty backwards. They guess what “hard enough” looks like, pick a cut score that feels appropriately tough, and then wonder why solid candidates keep getting filtered out. The fix is unglamorous. Get three or four of your best performers in that role to take the test blind, with no idea it will ever be used against anyone. Their score range becomes your real floor instead of a number someone picked because it sounded selective. If your best engineer would have failed the cut score you just set, the test is broken, not the engineer.

Most false negatives come down to the same mistake. A rigid cut score treats “didn’t know this specific term” the same as “couldn’t reason through the problem,” and those aren’t remotely the same failure. So instead of one hard line, I look at what kind of mistake put someone just under it. A wrong answer backed by sound reasoning gets a short follow-up conversation before anyone says no, not an automatic reject. That’s fifteen extra minutes per borderline candidate, and in most cases that short conversation is what saves a good hire the score alone would have thrown away.

How do you connect pre-hire assessment data to post-hire performance to validate and improve your hiring funnel within a 90-day window?

I don’t build a new review process to make this work.

I ask the hiring manager for two or three concrete things they’d already say about someone at the 90-day mark, in their own words. Examples include “hasn’t needed hand-holding on deployments” or “still checks with me before touching the billing code.”

Those are real observations a manager forms anyway. I just make sure they get written down and tied back to that person’s assessment sub-scores instead of living only in the manager’s head.

The useful part isn’t where the score and the manager’s read agree. It’s where they don’t.

If someone scored high on a specific section and the manager’s 90-day read says they’re still struggling with that exact thing, that section is telling you less than you think. You should find out why before you rely on it for the next candidate.

Do that across a handful of hires in a role, and the weighting on your assessment starts moving toward what the job actually rewards, not what looked impressive on a test three months earlier.

It’s slow; it’s a small sample every time, and it beats waiting a year for a performance review cycle to tell you the same thing.

When offering free assessments or highly accessible screening, what safeguards do you use to ensure fairness, prevent gaming, and preserve a positive candidate experience?

Fairness starts with everyone getting the same test at the same difficulty in the same time window. It sounds obvious, but many “free assessment” tools quietly serve different question sets to different batches of candidates, and then you’re not comparing people, you’re comparing test versions. We rotate the question bank on a schedule instead of on the fly, so a candidate who takes it Monday and a candidate who takes it Thursday aren’t accidentally getting an easier or harder draw.

Gaming mostly shows up as a pattern, not a single suspicious answer. We flag tests for review when we see patterns such as:

  • a cluster of candidates finishing in under two minutes
  • the same exact wrong answer showing up across people who’ve never met

Either of those gets a test flagged for review before it gets flagged for anyone’s individual score.

The candidate experience part is the one people skip when the assessment is free, and it’s the one that actually protects the brand. Every candidate gets a real report back, not just the ones who pass, even if it’s short. If we tell someone it takes fifteen minutes, it takes fifteen minutes, because nothing burns trust faster than a “quick screener” that eats forty-five minutes of someone’s evening for a job they don’t get. None of these are exotic controls; they’re just the ones that get cut first when a team decides free means low effort, and that’s usually the mistake.

Thanks for sharing your knowledge and expertise. Is there anything else you'd like to add?

Just one thing: none of what we talked about replaces a hiring manager’s judgment. It provides better inputs to apply that judgment.

The teams actually getting hiring right this year aren’t the ones with the fanciest assessment stack; they’re the ones who read the report before deciding instead of just checking whether a candidate cleared some number. If you take one thing from this, skip the score for a minute and read the actual answer a candidate gave—that’s usually where the real information is.

Up Next