Senior AI SDET · AI Evaluation · Agent Testing · Red-Teaming · open to global remote
I am _
AI ships fast and lies confidently. I've spent 8 years building the
tests that catch it — hallucinating models, flaky agents,
green checks that mean nothing — before a single user ever sees it.
Everyone's shipping AI. Someone has to ask it the hard questions first —
that's me.
For 8 years I've worked at the intersection of AI and
quality engineering. Today I'm a Senior AI SDET at Sentient
Labs, leading AI Quality & Evaluation for AGI-powered
agents: hallucination detection, reasoning validation, tool-use
accuracy, prompt adherence, gold datasets, and the regression
pipelines that turn "it seems fine" into
measured confidence — plus agentic and
crypto/Web3 workflows, from wallet integrations to on-chain checks.
Before this, I was the founding QA engineer at Level AI,
building the quality function from zero — voice AI, conversational AI,
STT accuracy, intent detection — back when testing AI meant inventing
the playbook as you went.
I attack the system before anyone else can — jailbreak & prompt-injection probes,
adversarial prompt suites, tool-abuse scenarios, data-leak attempts, failure injection.
A model that's only been asked nice questions isn't tested; it's flattered.
0%
03. work
Jul 2025 — now · Bangalore
Senior AI SDET
Sentient Labs — AGI startup, autonomous AI agents
Leading AI Quality & Evaluation for AGI-powered agents. Built eval
frameworks for hallucination detection, reasoning validation, tool-use
accuracy, and prompt adherence. Created gold datasets & AI regression
pipelines — improving release confidence by ~40% and cutting flaky
validations by ~30%. Playwright web automation and REST Assured /
Postman API frameworks from scratch; tested agentic + crypto/Web3
workflows including wallets and on-chain / off-chain data checks.
llm evals
agent testing
gold datasets
web3
playwright
Oct 2021 — Jul 2025 · New Delhi
Senior SDET · Founding Engineer
Level AI — voice & conversational AI
Founding QA hire: built the QA function 0 → 1 — automation strategy,
release gates, quality processes. Scalable UI + API frameworks
(Selenium, Playwright, REST Assured) improving regression coverage by
~70%, integrated into CI/CD with BDD. Led voice & conversational AI
testing — STT accuracy, intent detection, entity extraction, call
summarization, agent assist. Performance testing with K6 / JMeter;
mentored SDETs.
0→1 qa org
voice ai
ci/cd
k6 / jmeter
mentoring
Sep 2020 — Oct 2021 · Pune
SDET-I
Motifworks
Automated 300+ hybrid web + desktop test scenarios with Selenium,
Java, C#, and SpecFlow — cutting regression effort by ~70%. Desktop
automation via FlaUI, mobile coverage with Appium.
selenium
specflow
flaui
appium
Jun 2018 — Aug 2020 · Pune
QA Automation Engineer
Atos
Automated banking web apps with Selenium, Protractor (TypeScript), and
Cucumber — BDD regression suites wired into CI pipelines.
banking
protractor
cucumber
bdd
2014 — 2018 · Bangalore
B.Tech, Computer Science & Engineering
Visvesvaraya Technological University
Where the habit of asking "but does it actually work?" began.
04. why hire me // reviewed like a pull request
⎇ candidate/deep-halder → your-team/main
hire: deep_halder #2026
✓ approved — strong hire
✓
Trusted twice as the first quality hire. Two AI startups handed me a blank page; both got a working QA org. That trust isn't given for buzzwords.
✓
Receipts, not vibes. ~40% lift in release confidence. ~30% fewer flaky validations. ~70% more regression coverage. I measure my own work the way I measure models.
✓
I catch what AI reviewers rubber-stamp. An LLM will confidently approve the bug it can't trace. My harnesses are built to be harder to fool than I am.
✓
"Not ready" is a complete sentence. Release gates only matter if someone is willing to hold them. I am — and I bring the data that ends the argument.