I build AI that has to work on Monday: vision-LLM graders that land within one mark of the teacher,
agents that chase failed payments without spamming anyone, and the unglamorous backend underneath.
Unfamiliar stack? I learn it on the way.
Fig. 00b / learning ledger[SINCE 2024]
every stack I picked up because something had to ship. 16 in two years, about one every six weeks.
20252026
2026-10 · 16 / 16
Postgres row-level security
because one school's fees are not another school's business. Kwilo admissions and fees.
Fig. 01 / shell[LIVE]
nishanth-lab shell 1.0. type , or tap a command.
try:
tab completes · ↑↓ history · esc leaves
location
Bengaluru, IN · UTC+5:30
availability
[OPEN TO: AI/ML ROLES, 2027]
current focus
Kwilo AI · federated IoT paper
last updated
§01 / work
Selected work
6 entries. each states the problem, the approach, the result, and what it cost me to learn.
Fig. 02 / project / 2026[SHIPPED]
Kwilo AI · production research
Vision-LLM exam grader
problemMarking handwritten university scripts eats teachers' evenings, and a model that is two marks off per subpart is worse than no model.
approachScanned PDF to page images, pages mapped to each subpart, then a vision LLM reads the marking scheme plus a few teacher-marked exemplars. Benchmarked GPT-5.4 Mini, Gemini 3 Flash and 3.1 Pro over 4 subjects and 11 test cells.
python
azure openai
vertex ai
pymupdf
uv
matplotlib
pooled MAE, marks per subpart (few-shot subparts)
0.855
error vs zero-shot, GPT (1.259 to 0.910)
−28%
subparts within ±1 mark of the teacher
81.7%
what I learnedExemplars beat model size: few-shot Flash (0.890) beat zero-shot Pro (1.195). And a harness that writes run manifests is worth more than any single prompt tweak, because it turns "feels better" into a number.
show the pipelineFig. 02a. One subpart per call in the subparts pipeline; that variant had zero script failures.
Fig. 03 / project / 2026[SHIPPED]
Razorpay AI Buildathon 2026 · track 03
PayRecover
problemA failed UPI or card payment can't be silently re-charged. Merchants either spam the customer or lose the sale.
approachRules diagnose the failure; Gemini only sees the ambiguous ones. Every action (wait, link, remind, escalate, stop) passes a policy gate with caps and a kill switch, and lands in an append-only audit log.
python
razorpay api
gemini
sqlite
pytest
of failed revenue recovered (INR 44,205 of 141,353)
31.3%
of recoverable customers captured
76.1%
seeded failure cases, one dry run
80
what I learnedThe model call is the easy 10%. The agent is the policy around it: typed timeouts, writes that can't double-fire, and a stop that stays stopped after you flip the switch back.
show the loopFig. 03a. The LLM sits inside "diagnose". It never gets to skip the gate.
problemSchools run exams, fees and admissions on this product, so a bug costs real money or leaks a real parent's phone number.
approachAnswer sheets auto-matched to students (filename first, then header OCR, never blocking the upload). A public blog stack end to end. Fee receipts with race-safe installment settlement. Admissions with a primary guardian and a one-page A4 form.
guardians per admission, one becomes the parent login
2
what I learnedA self-reported phone number is not proof of identity, so admitting a student now asks staff before linking a parent account. And lock the row before two payments settle the same installment milliseconds apart.
Research project · IoT security · paper on the way
Federated IoT anomaly detector
problemIoT networks need intrusion detection, but shipping raw traffic from every camera and sensor to one server is a privacy and bandwidth problem.
approachDockerised federated clients for cameras, sensors and controllers; RFRE preprocessing on IP and port features; FedAvg vs FedProx aggregation benchmarked for stability against spoofing and DDoS.
python
pytorch
docker
fedprox
scikit-learn
aggregated AUC with FedProx
0.98
device classes simulated
3
raw packets leaving a client
0
what I learnedMost of what I called "model instability" was a data-partitioning bug. Check the shards before you tune the optimizer.
problemWhen a coastal hazard hits, the network is often the first thing to go, and the reports that do get out are a mix of real sightings, panic and recycled photos.
approachAn offline-first app that relays reports over a peer-to-peer BLE mesh, an AI scraper that reads social media, and a dashboard that clusters reports on a PostGIS map. My part: the IndicBERT sentiment pipeline over scraped posts and the image verification module that gives each report a confidence score.
python
indicbert
pytorch
hugging face
fastapi
supabase + postgis
react native
internet needed to file a report (BLE mesh relay)
0
pillars: mobile app, AI scraper, web dashboard
3
AI modules I owned: sentiment, image verification
2
what I learnedCoastal posts mix English, Hindi and regional languages in a single sentence, so an English-only model was never an option. And a report needs a confidence score before it reaches a map that people act on.
show the pipelineFig. 06a. Both signals feed one confidence score before a report lands on the map.
problemClaims spread faster than anyone can check who started them or who is amplifying them.
approachA LangGraph state machine: a collector picks a tool and gathers evidence, a verifier scores each item, a graph node maps authors and mentions, and a refiner decides whether to dig deeper or stop.
python
langgraph
pydantic
fastapi
next.js
graph nodes in the loop
5
trust score per evidence item
0–1
apps: agent API + Next.js UI
2
what I learnedAn agent loop needs an explicit stopping rule and a typed shared state before it needs a better model. Pydantic on every edge is what kept the demo alive.
show the graphFig. 07a. The refiner is the only node allowed to end the loop.
smaller experiments, hackathons and write-ups. including the ones that didn't work.
Fig. 08 / notebook.log[APPEND-ONLY]
Lab log of experiments and builds, newest first
date
entry
tag
result
2026-10
w2v-BERT 2.0 vs Whisper-large encoders
ML
OOF 0.9346, best so far. public LB didn't move; the changes landed in the private split.
2026-09
Fine-tuning Whisper-large on 826 clips
ML
lost to a frozen probe, 0.904 vs 0.925. overfits by epoch 5.
2026-09
Face-tracking water gun
IOT
MediaPipe face centre to pan/tilt angles to two ESP32 servos. shoot is still a stub.
2026-07
AssetFlow, 8-hour hackathon
HACK
owned the asset track: registry, allocation state machine, transfers, overdue flagger.
2026-05
Gemini 3.1 Pro vs 3 Flash, zero-shot grading
EVAL
Flash wins File Structures by 0.42 MAE, loses Math by 0.54. no free lunch.
2026-05
Sarvam document OCR as a grading front-end
EVAL
benchmarked, then removed. documented why so nobody re-tries it blind.
2026-04
Kannada name generator, char-level GRU
ML
loss 2.08 to 1.41 over 10k epochs. some names are real, most are merely plausible.
2025-12
OWASP Top 10 lab series
SEC
17/17 PortSwigger labs written up. the README updates itself via Actions.
2025-11
DFA-based intrusion detection, from scratch
SEC
Snort-style rules compiled to a DFA plus PCRE, TCP reassembly, live dashboard.
2025-11
Zero-knowledge medical records vault
SEC
AES-256-GCM with RSA-OAEP key wrap in the browser. the server only ever sees ciphertext.
2025-09
Samudra Prahari, Smart India Hackathon
HACK
offline BLE-mesh hazard reports with image verification before broadcast.
2025-08
REINFOSEC CTF sprint
SEC
100+ challenges in 8 weeks. TryHackMe top 2%.
2025-05
Green Quest, Social Hackathon
HACK
adopt a plant on a map, AI-verified care quests, and a plant you can chat with.
2024-10
Green Terrace, BuzzOnEarth @ IIT Kanpur
HACK
3rd place. first podium, first all-nighter that paid off.
§03 / about
About
the readable version of the resume.
about.md · utf-8 · lf[README]
I'm Nishanth. I build LLM systems that have to be right, then the boring code that keeps them right.
Final-year computer science at Bangalore Institute of Technology (IoT, cybersecurity and blockchain track, CGPA 8.50). AI/ML intern at Kwilo AI since April 2026, with a two-month break in the middle, working on answer-script grading and the product around it.
## how I work
I don't pick a lane. I pick a problem and learn whatever lane it lives in. In two years that has meant:
100+ CTFs to learn how web apps break (REINFOSEC, 2025)
Spring Boot to learn how they're built (Goloka IT, 2025–26)
LangGraph, to make an agent investigate instead of autocomplete
vision LLMs and an eval harness, to grade handwriting against teachers
Soroban, to store a notepad on the Stellar chain, mostly to see if I could
servo trigonometry, to aim a water gun at faces (it misses less now)
Postgres row-level security, because one school's fees are not another school's business
## currently
shipping ERP and AI features at Kwilo
squeezing emotion out of 826 Hindi clips before 12 Oct
290 LeetCode problems in, 253 of them in Java
## off the clock
Lazy Monks: four of us, 10+ hackathons, one podium (3rd at BuzzOnEarth, IIT Kanpur). First place at Turing Much, the college ML competition. Core member of cryptX, the college technical team.