N.A.lab notebookvol. 2026, Nishanth Antony, back to top

Fig. 00 / subject  ·  AI/ML engineer, Bengaluru

Nishanth Antony.

I build AI that has to work on Monday: vision-LLM graders that land within one mark of the teacher, agents that chase failed payments without spamming anyone, and the unglamorous backend underneath. Unfamiliar stack? I learn it on the way.

Fig. 00b / learning ledger[SINCE 2024]

every stack I picked up because something had to ship. 16 in two years, about one every six weeks.

2026-10 · 16 / 16

Postgres row-level security

because one school's fees are not another school's business. Kwilo admissions and fees.

Fig. 01 / shell [LIVE]
nishanth-lab shell 1.0. type , or tap a command.
try:

tab completes · ↑↓ history · esc leaves

location
Bengaluru, IN · UTC+5:30
availability
[OPEN TO: AI/ML ROLES, 2027]
current focus
Kwilo AI · federated IoT paper
last updated
§01 / work

Selected work

6 entries. each states the problem, the approach, the result, and what it cost me to learn.

Fig. 02 / project / 2026[SHIPPED]

Kwilo AI · production research

Vision-LLM exam grader

problemMarking handwritten university scripts eats teachers' evenings, and a model that is two marks off per subpart is worse than no model.

approachScanned PDF to page images, pages mapped to each subpart, then a vision LLM reads the marking scheme plus a few teacher-marked exemplars. Benchmarked GPT-5.4 Mini, Gemini 3 Flash and 3.1 Pro over 4 subjects and 11 test cells.

  • python
  • azure openai
  • vertex ai
  • pymupdf
  • uv
  • matplotlib
pooled MAE, marks per subpart (few-shot subparts)
0.855
error vs zero-shot, GPT (1.259 to 0.910)
−28%
subparts within ±1 mark of the teacher
81.7%
what I learnedExemplars beat model size: few-shot Flash (0.890) beat zero-shot Pro (1.195). And a harness that writes run manifests is worth more than any single prompt tweak, because it turns "feels better" into a number.
show the pipeline
Pipeline: scanned PDF, page images, subpart page map, prompt builder, vision LLM, marks with evidence, metrics against held-out teacher marks scanned PDFanswer script page imagesrendered subpart map2a -> pp. 2-4 promptrubric + exemplars vision LLMgpt / gemini marks + evidenceJSON per subpart metricsMAE, ±1, QWK teacher marksground truth dashed: never shown to the model, used only for scoring
Fig. 02a. One subpart per call in the subparts pipeline; that variant had zero script failures.
Fig. 03 / project / 2026[SHIPPED]

Razorpay AI Buildathon 2026 · track 03

PayRecover

problemA failed UPI or card payment can't be silently re-charged. Merchants either spam the customer or lose the sale.

approachRules diagnose the failure; Gemini only sees the ambiguous ones. Every action (wait, link, remind, escalate, stop) passes a policy gate with caps and a kill switch, and lands in an append-only audit log.

  • python
  • razorpay api
  • gemini
  • sqlite
  • pytest
of failed revenue recovered (INR 44,205 of 141,353)
31.3%
of recoverable customers captured
76.1%
seeded failure cases, one dry run
80
what I learnedThe model call is the easy 10%. The agent is the policy around it: typed timeouts, writes that can't double-fire, and a stop that stays stopped after you flip the switch back.
show the loop
Loop: failed payment, diagnose, policy gate, act, all steps written to an append-only audit log; act loops back to diagnose on the next tick next tick failed paymentrazorpay diagnoserules, then LLM policy gatecaps, kill switch actlink / remind / stop audit log: append-only, one corr id per case
Fig. 03a. The LLM sits inside "diagnose". It never gets to skip the gate.
Fig. 04 / project / 2026[IN PROD]

Kwilo AI · Apr–May, Jul 2026–now

Shipping on an AI learning platform

problemSchools run exams, fees and admissions on this product, so a bug costs real money or leaks a real parent's phone number.

approachAnswer sheets auto-matched to students (filename first, then header OCR, never blocking the upload). A public blog stack end to end. Fee receipts with race-safe installment settlement. Admissions with a primary guardian and a one-page A4 form.

  • fastapi
  • sqlalchemy async
  • postgres + RLS
  • alembic
  • react 19
  • typescript
  • tanstack query
  • gemini
uploads blocked when OCR can't read a sheet
0
product areas: exams, content, blog, fees, admissions
5
guardians per admission, one becomes the parent login
2
what I learnedA self-reported phone number is not proof of identity, so admitting a student now asks staff before linking a parent account. And lock the row before two payments settle the same installment milliseconds apart.
Fig. 05 / research / 2026[RESEARCH]

Research project · IoT security · paper on the way

Federated IoT anomaly detector

problemIoT networks need intrusion detection, but shipping raw traffic from every camera and sensor to one server is a privacy and bandwidth problem.

approachDockerised federated clients for cameras, sensors and controllers; RFRE preprocessing on IP and port features; FedAvg vs FedProx aggregation benchmarked for stability against spoofing and DDoS.

  • python
  • pytorch
  • docker
  • fedprox
  • scikit-learn
aggregated AUC with FedProx
0.98
device classes simulated
3
raw packets leaving a client
0
what I learnedMost of what I called "model instability" was a data-partitioning bug. Check the shards before you tune the optimizer.
Fig. 06 / hackathon / 2025[SIH 2025]

Smart India Hackathon 2025 · INCOIS · PS ID25039

Samudra Prahari

problemWhen a coastal hazard hits, the network is often the first thing to go, and the reports that do get out are a mix of real sightings, panic and recycled photos.

approachAn offline-first app that relays reports over a peer-to-peer BLE mesh, an AI scraper that reads social media, and a dashboard that clusters reports on a PostGIS map. My part: the IndicBERT sentiment pipeline over scraped posts and the image verification module that gives each report a confidence score.

  • python
  • indicbert
  • pytorch
  • hugging face
  • fastapi
  • supabase + postgis
  • react native
internet needed to file a report (BLE mesh relay)
0
pillars: mobile app, AI scraper, web dashboard
3
AI modules I owned: sentiment, image verification
2
what I learnedCoastal posts mix English, Hindi and regional languages in a single sentence, so an English-only model was never an option. And a report needs a confidence score before it reaches a map that people act on.
show the pipeline
Pipeline: scraped social posts go through IndicBERT sentiment, field reports with photos go through image verification, both feed a confidence score per report, which lands on the PostGIS dashboard social postsweb scraper field reportphoto, BLE mesh IndicBERTsentiment image checkverify the photo confidencescore per report dashboardPostGIS map orange: the two modules I built. the rest is my teammates' work
Fig. 06a. Both signals feed one confidence score before a report lands on the map.
Fig. 07 / project / 2025[HACKATHON]

Recurzive V2 · Lazy Monks

Veris truth engine

problemClaims spread faster than anyone can check who started them or who is amplifying them.

approachA LangGraph state machine: a collector picks a tool and gathers evidence, a verifier scores each item, a graph node maps authors and mentions, and a refiner decides whether to dig deeper or stop.

  • python
  • langgraph
  • pydantic
  • fastapi
  • next.js
graph nodes in the loop
5
trust score per evidence item
0–1
apps: agent API + Next.js UI
2
what I learnedAn agent loop needs an explicit stopping rule and a typed shared state before it needs a better model. Pydantic on every edge is what kept the demo alive.
show the graph
Graph: query to collector, verifier, graph builder, refiner; refiner loops back to collector with a new query or ends at the reporter next_query query collectorpicks a tool verifiertrust 0-1 graphwho said it refinerdig / stop reporterdone
Fig. 07a. The refiner is the only node allowed to end the loop.
§02 / notebook

Lab log

smaller experiments, hackathons and write-ups. including the ones that didn't work.

Fig. 08 / notebook.log[APPEND-ONLY]
Lab log of experiments and builds, newest first
dateentrytagresult
2026-10w2v-BERT 2.0 vs Whisper-large encodersMLOOF 0.9346, best so far. public LB didn't move; the changes landed in the private split.
2026-09Fine-tuning Whisper-large on 826 clipsMLlost to a frozen probe, 0.904 vs 0.925. overfits by epoch 5.
2026-09Face-tracking water gunIOTMediaPipe face centre to pan/tilt angles to two ESP32 servos. shoot is still a stub.
2026-07AssetFlow, 8-hour hackathonHACKowned the asset track: registry, allocation state machine, transfers, overdue flagger.
2026-05Gemini 3.1 Pro vs 3 Flash, zero-shot gradingEVALFlash wins File Structures by 0.42 MAE, loses Math by 0.54. no free lunch.
2026-05Sarvam document OCR as a grading front-endEVALbenchmarked, then removed. documented why so nobody re-tries it blind.
2026-04Kannada name generator, char-level GRUMLloss 2.08 to 1.41 over 10k epochs. some names are real, most are merely plausible.
2025-12OWASP Top 10 lab seriesSEC17/17 PortSwigger labs written up. the README updates itself via Actions.
2025-11DFA-based intrusion detection, from scratchSECSnort-style rules compiled to a DFA plus PCRE, TCP reassembly, live dashboard.
2025-11Zero-knowledge medical records vaultSECAES-256-GCM with RSA-OAEP key wrap in the browser. the server only ever sees ciphertext.
2025-09Samudra Prahari, Smart India HackathonHACKoffline BLE-mesh hazard reports with image verification before broadcast.
2025-08REINFOSEC CTF sprintSEC100+ challenges in 8 weeks. TryHackMe top 2%.
2025-05Green Quest, Social HackathonHACKadopt a plant on a map, AI-verified care quests, and a plant you can chat with.
2024-10Green Terrace, BuzzOnEarth @ IIT KanpurHACK3rd place. first podium, first all-nighter that paid off.
§03 / about

About

the readable version of the resume.

about.md · utf-8 · lf[README]

I'm Nishanth. I build LLM systems that have to be right, then the boring code that keeps them right.

Final-year computer science at Bangalore Institute of Technology (IoT, cybersecurity and blockchain track, CGPA 8.50). AI/ML intern at Kwilo AI since April 2026, with a two-month break in the middle, working on answer-script grading and the product around it.

## how I work

I don't pick a lane. I pick a problem and learn whatever lane it lives in. In two years that has meant:

  • 100+ CTFs to learn how web apps break (REINFOSEC, 2025)
  • Spring Boot to learn how they're built (Goloka IT, 2025–26)
  • LangGraph, to make an agent investigate instead of autocomplete
  • vision LLMs and an eval harness, to grade handwriting against teachers
  • Soroban, to store a notepad on the Stellar chain, mostly to see if I could
  • servo trigonometry, to aim a water gun at faces (it misses less now)
  • Postgres row-level security, because one school's fees are not another school's business

## currently

  • shipping ERP and AI features at Kwilo
  • squeezing emotion out of 826 Hindi clips before 12 Oct
  • 290 LeetCode problems in, 253 of them in Java

## off the clock

Lazy Monks: four of us, 10+ hackathons, one podium (3rd at BuzzOnEarth, IIT Kanpur). First place at Turing Much, the college ML competition. Core member of cryptX, the college technical team.

Fig. A / subject[NO PHOTO]
photo: N/A. which are also my initials.
Fig. B / facts[VERIFIED]
based in
Bengaluru
graduating
Jul 2027
CGPA
8.50 / 10
LeetCode
290
TryHackMe
top 2%
certs
OCI GenAI Pro, APIsec ACP
§04 / skills

Skills, honestly

no logo wall. you can't grep a logo.

Fig. 09 / skills.txt[SELF-REPORTED]
Skills by domain and how often I use them
domain dailyused this week comfortableshipped it, no docs for basics dabblingshipped once, docs open
LLMs & agents
  • few-shot prompt design
  • Gemini on Vertex AI
  • Azure OpenAI
  • eval harnesses
  • LangGraph
  • LangChain
  • RAG with pgvector
  • AWS Bedrock
  • multi-agent loops
  • Sarvam AI
  • Claude API
Models & data
  • pandas
  • NumPy
  • PyTorch
  • scikit-learn
  • Whisper / w2v-BERT probing
  • matplotlib
  • TensorFlow
  • OpenCV
  • MediaPipe
  • federated learning
Backend
  • FastAPI
  • SQLAlchemy async
  • PostgreSQL
  • Alembic
  • pytest
  • Flask
  • Spring Boot
  • REST
  • WebSockets
  • Postgres RLS
  • Supabase
  • Firebase
  • Prisma
Frontend
  • React
  • TanStack Query
  • Tailwind
  • Vite
  • React Native (Expo)
  • vitest
  • Next.js
  • Phaser
MLOps & cloud
  • Git
  • Docker
  • CI on GitHub Actions
  • GCP Cloud Run + GCS
  • Kaggle GPU pipelines
  • uv
  • Azure Container Apps
  • AWS EC2 / S3
  • Cloudflare
Languages
  • Python
  • TypeScript
  • Java
  • SQL
  • Bash
  • C
  • C++
  • Rust (Soroban)
Security none right now (was daily in 2025)
  • Burp Suite
  • nmap
  • OWASP Top 10
  • API security
  • gobuster
  • sqlmap
  • hydra
  • hashcat
  • Scapy
  • Web Crypto
Fig. 10 / contact[OPEN]

Got a problem that has to work?

leetcode
u/Nish345
$ open mailto:nishanthantony5@gmail.com