A local Vitest run against the repository working tree on 2026-08-28; external services are mocked at their test boundaries.
AI engineer / Bengaluru / building end to end
I build AI products, break them, fix them, and write down what happened.
I’m Parth, an AI engineer in Bengaluru. This is the work, the evidence behind it, and the part that broke.

The case studies are open now. Published worlds open first, then hand off to the engineering record.
Enter BeatMind ↗Stem-separation benchmark for one fixed 120-second audio input; this is not full pipeline latency.
The correction recovered 4 of 7 initial failures, a 5.7 percentage-point lift; 12 adversarial queries are a separate set.
The person on the sheet
Hi, I am Parth.
I am Parth, an AI engineer in Bengaluru. I like products that feel alive on the surface and stay strict underneath.
The work here includes music tools, retrieval systems, agents, diffusion pipelines, fraud models, and small automations. Every case study includes the part that failed.
More about how I work →
The register
Work, in the order I would show it.
Every project stays on the sheet. The full register will sort by build effort, recency and current status.
- 01BeatMindLive · Flagship↗
A music workspace that separates a song, understands its structure, lets people rearrange it, and survives failed workers.
- 02VividLive · Flagship↗
A script-to-storyboard system that plans a scene as a sequence and tries to keep the same people recognisable from shot to shot.
- 03TathyaIn progress · Flagship↗
A sourced public record that groups what institutions and publishers said without turning the system into a judge.
- 04MedRAGShipped · Substantial↗
A drug-information retrieval system designed to cite the evidence it has and refuse questions it cannot support.
- 05Order SupervisorShipped · Focused↗
A durable order workflow where the model can propose actions but never becomes the source of truth for the order lifecycle.
- 06QueryPilotShipped · Substantial↗
A natural-language-to-SQL API that retrieves schema context, validates generated SQL, and gives failed queries one bounded correction loop.
- 07SecondSelfRunning · Flagship↗
An evidence-bound career system that prepares applications and stops at a human review queue before anything consequential is sent.
- 08OncoVerseIn progress · Substantial↗
A cancer education atlas that makes anatomy and disease progression visible while keeping every explanation inside a source boundary.
- 09UPI Fraud EngineShipped · Substantial↗
A real-time fraud scoring system evaluated at a fixed alert budget instead of optimising a model metric in isolation.
- 10Spur ChatTake-home · Focused↗
A small streaming support assistant built to a take-home brief and bounded to one fictional brand's catalogue and policies.
- 11Fraud Risk IntelligenceShipped · Focused—
An earlier fraud modelling system where training, serving, and explanations share one frozen preprocessing contract.
- 12Oracle Auto ProvisionRunning · Focused—
A small scheduled utility that retries scarce Oracle Cloud capacity without creating a duplicate instance.
Measured, with the denominator attached
Proof without the victory lap.
The current BeatMind web workspace passes 381 tests across 39 test files.
Verified 2026-08-28 / 39 passing test filesOn the documented 120-second benchmark input, BeatMind's L4 separation run took 56.5 seconds versus 97.2 seconds on T4.
Verified 2026-08-28 / one 120-second benchmark track on each GPU profileQueryPilot's correction loop moved execution success from 63 to 67 queries on the 70-query core set.
Verified 2026-08-28 / 70 core benchmark queriesThree kinds of work
Where I can be useful.
AI products people can use
A working product in the browser, with the model, data, evaluation, and interface connected.
- Retrieval over private material
- Agents with bounded actions
- Generation workflows with durable jobs
Business workflow automation
Repeated work moves through a visible, recoverable workflow while people keep authority over consequential decisions.
- Long-running workflow orchestration
- Risk and fraud scoring
- Scheduled infrastructure tasks
Interactive product storytelling
A web experience that explains how the system works without hiding the useful page behind motion.
- Data-led visual stories
- Accessible interaction and reduced motion
- Static fallbacks that keep the whole argument
Errata and writing
The parts I got wrong stay public.
The consent box I should not have removed
A cleaner upload flow removed a constraint the product was responsible for keeping.
The explanation needs the same input
Separate transformations could make the API score and the explanation disagree.
A refusal is part of the MedRAG result
Fluent answers initially received more attention than evidence sufficiency.
The client door, again
Have something real to build?
Tell me what you are trying to make. I will tell you what I can own, what needs proving first, and the smallest useful way to start.





