AI Lab
AI Systems · Product Architecture · Evaluation
I design AI products and production systems around judgment, reliability, evaluation and human control — not just task automation.
OPERATIONAL AI AT SCALE
A production AI pipeline that processed 173,517 company records end-to-end in 70 hours, combining concurrent processing, confidence scoring, deterministic validation and human-in-the-loop review.
The core challenge was not calling an LLM at scale. It was designing around quota limits, uncertainty, cost, failure recovery and the boundary between probabilistic judgment and deterministic control.
Read the full case study173,517
Companies processed
99.996%
Processing success
70 hrs
End-to-end runtime
$216.50
Total API cost
BEHAVIORAL MEMORY SYSTEM
A behavioral memory system that learns durable user preferences without turning every correction into permanent behavior.
The dangerous memory error is inventing permanence the user never intended.
7/8
Machine end-to-end
8/8
Human-adjudicated
96.7%
Memory action accuracy
GENERATIVE AI SYSTEM
Collaboratively designed a production generative workflow that separates creative generation from deterministic controls, then turns human editorial feedback into governed context and reusable rules.
The harder problem was not generating drafts. It was keeping context, memory and human authority reliable as the system evolved.
31+
Live articles
8–12
Posts / week capacity
≤3
Writer–critic iterations
AI-ASSISTED 0→1 PRODUCT
A Chrome productivity product built independently from idea to launch. The core product challenge was not the code — it was choosing a narrow user segment, designing around reading backlog, and finding a monetisation boundary that protected the core habit.
AI accelerated execution, but segmentation, monetisation and distribution remained product decisions.
140+
Chrome Web Store installs
0
Paid acquisition
v1.1.0
Current release
₹799/year
Pro Yearly pricing
01 · Judgment
Use language models where interpretation is needed. Use deterministic systems for rules, validation and schema enforcement.
Design for probabilistic output with confidence thresholds, fallback paths and graceful degradation.
02 · Reliability
Define what 'good' means before tuning prompts, adding agents or scaling the workflow.
Logging, tracing and confidence visibility turn silent failures into actionable product feedback.
03 · Control
Human-in-the-loop is a design decision, not a failure state. Build review and escalation into the workflow.
Use evaluation results, human corrections and structured outputs to improve future decisions without implying autonomous self-learning.
Responsibility
Systems
Economics
Active areas of exploration — not claims of expertise.
05 · Closing
AI products are still products. But increasing capability raises the bar for how deliberately they are evaluated, governed and deployed.
The goal is not just more capable systems — it is systems that are useful, observable and responsibly deployable.