KRISH ARYA — SYSTEM

AI Engineer · Builder · Researcher

KRISH ARYAI build
intelligent
systems.

Production voice, retrieval and evaluation at EaseMyTrip. A product shipped alone at Tov AI. Published research on systems that correct themselves.

System mapResearch feeds AI models, which branch into RAG, agents and voice. Those meet in systems, which split into products and evaluation, and both flow into the engineering workflow. Evaluation feeds back into the models. Each project hangs off the capability it is built on.RESEARCHAI MODELSRAGAGENTSVOICESYSTEMSPRODUCTSEVALUATIONENG. WORKFLOWNL2SQLSCAM DETECTIONXAIENV. MLSUPPORT RAGEASYDARSHANVOICE AGENTV2V AGENTBYLIN.APPMUTAGENLAUNCHMYFOLIOEVAL · 50K+MIRA PILOT
[ Enter the system ]
Voice agent, end to end
<200ms
Conversations scored
50K+
Spider accuracy, published
92%
Users in week one
1,000+

00 · Index

The model is one box.The engineering is the loop around it.

Almost every system below has the same shape: something generates, something real verifies, and the result is fed back. A test has to kill a mutant. SQL has to execute. A claim has to exist in its source. A finding has to survive an engineer.

Generate, verify and feed back, for each system
System01 Generates02 Verified by03 Feeds back
B-02MutagenLLM writes a testDoes it kill the injected mutant?Survivors are regenerated
EXP-001NL2SQLGemini writes SQLExecution against the databaseFailure + schema + error → vector memory
E-02Support RAGChat logs → Q/A recordsPII scrub, contradiction check vs KBApproved records re-embedded
B-01Bylin.appMulti-model generationEmbedding similarity + LLM gradeApproval history biases later outputs
E-04EasyDarshanSchema-constrained extractionIs the claim in the source snapshot?Unsupported claims dropped
E-03Eval platformThe support botComposite bot-health scoreFour channels show where it fails
E-05V2V agentSpeculative prefill on a partialDid the partial change?Discard and reissue
I-01Mira pilotAI first-pass reviewEngineer accepts or ignoresSignal quality decides adoption

World 01 · Shipping products

Build

Products that went from an idea to people using them. One as the sole engineer at a company, one as open source, one with a thousand users in its first week.

B-01Tov AIAI EngineerApr–Jun 2026Remote

Bylin.app

An AI content-generation product. Idea to live users, one engineer.

L1

An AI content-generation product built at Tov AI and taken from idea to live users. I was the sole engineer: model pipeline, database, deployment and instrumentation.

L2

Every request runs a multi-stage pipeline: summarization, retrieval, generation, quality scoring. It runs across multiple models via OpenRouter and is orchestrated on self-hosted n8n. Users then approve or reject what comes out.

Engineer (sole)
1
Pipeline stages
4
Shipped to users
Live
  • OpenRouter
  • n8n (self-hosted)
  • Postgres
  • Supabase
  • Embeddings
  • LLM grading

--/09Generation is one stage of four. Scoring decides what ships.

B-02Open sourcePythonpip install mutagen-ai

Mutagen

LLM-driven mutation testing. A test only survives if it catches a planted bug.

L1

LLM-driven test generation where every generated test must kill an injected mutant to survive.

L2

An LLM writes a test. Mutagen injects a mutant, a small deliberate fault in the code under test. The test runs against the mutant. If the test fails, it detected the fault: the mutant is killed and the test is kept. If the test still passes, it can't see the fault, so it gets regenerated.

Lines of code
~11K
Success criterion, not coverage
Kill
mutagen-ai
PyPI
  • Python
  • LLMs
  • Mutation testing
  • FSM
  • asyncio
  • Hexagonal architecture

--/11Success ≠ coverage. Success = demonstrated fault detection.

Payload appears here as the loop runs

Example function and tests are illustrative. The loop is Mutagen's.

B-03ProductGroqFlaskNext.js

LaunchMyFolio

A resume goes in. A personalized portfolio site comes out.

L1

An AI platform that turns a resume into a personalized portfolio site, end to end. 1,000+ users in its first week.

L2

The resume is read and understood, then the system generates tailored copy, picks the sections and lays them out for each user.

Users in week one
1,000+
Things generated: copy, sections, layout
3
  • Groq
  • Flask
  • Next.js

--/05The model's output is an interface, not a paragraph.

Payload appears here as the loop runs

Payloads are schematic.

World 02 · Complex AI systems in production

Engineer

Systems running against real traffic, with real latency budgets and real failure modes.

EaseMyTrip — production

Technology Trainee (AI/ML) · Jun 2026 – present · New Delhi

E-01EaseMyTripProductionLiveKit

Real-time voice agent

In-house outbound calling. Speech in, speech out, under 200ms.

L1

An in-house real-time voice agent for outbound calling, with under 200ms end-to-end latency.

L2

STT, LLM and TTS run as one streaming pipeline instead of three calls in sequence. Audio streams into speech-to-text, text streams into the model, and tokens stream into text-to-speech. LiveKit handles transport.

End-to-end latency
<200ms
Streaming stages, overlapped
3
  • STT
  • LLM
  • TTS
  • LiveKit
  • Streaming
Illustrative timing, not to scale
USER AUDIOSTREAMING STTLLMSTREAMING TTSAUDIO OUT<200MS END TO END

Stages overlap: each starts on the previous one's partial output instead of waiting for it to finish.

E-02EaseMyTripProductionpgvector

Production RAG

A customer-support knowledge base that updates itself from real conversations.

L1

Took the customer-support chatbot to a production RAG system whose knowledge base keeps itself current.

L2

Scattered support content was consolidated into one deduplicated knowledge base on a Dockerized pgvector store.

A pipeline then turns raw chat logs into (intent, question, answer, confidence) records. It scrubs PII, checks each record against the KB for contradictions, and re-embeds the additions that are approved.

Knowledge base
Self-updating
Fields per learned record
4
  • pgvector
  • Postgres
  • Docker
  • Embeddings
  • PII scrubbing

--/11Two loops: one answers, one learns.

Payload appears here as the loop runs

Example conversation and record are illustrative.

E-03EaseMyTripFastAPIReact

Evaluation platform

Defined how the support bot is measured. Then built the instrument.

L1

Defined the chatbot's evaluation framework from scratch and built the platform that runs it, scoring 50K+ conversations in parallel.

L2

A composite bot-health score combines four signals: escalations, tool-call success, intent resolution and friction. A FastAPI service scores conversations in parallel, with a React front end.

Conversations scored
50K+
Health channels
4
  • FastAPI
  • React
  • Python
  • Eval pipelines

Conversations scored

50K+

bot health = f(escalations, tool success, resolution, friction)

CH1 ESCALATIONShanded to a humanCH2 TOOL-CALL SUCCESSdid the action workCH3 INTENT RESOLUTIONwas the need metCH4 FRICTIONhow hard it wasBOT HEALTHcomposite

Click a channel to isolate it. Four signals separate the ways a conversation fails; the composite says how much.

Signal shapes are illustrative, not production data.

E-04EaseMyTripAgenticGoing live

EasyDarshan

An agentic pilgrimage planner that decides from evidence, not from fluency.

L1

An agentic AI planner for multi-stop pilgrimage itineraries under real constraints. I designed the agent loop and shipped the full-stack app around it.

L2

The agent plans through tool calls over retrieval and computed data instead of single-shot generation. It checks seasonal closures, mandatory registration, travel feasibility and accessibility.

The app around it has admin roles and a human review queue.

Entities discovered via SPARQL
2,000+
Hard constraints per plan
4
  • Agents
  • Tool calling
  • SearXNG
  • Tavily
  • SPARQL
  • Scrapers
  • Structured extraction

--/09The agent decides from evidence. Unsupported claims never reach the plan.

Payload appears here as the loop runs

Claims shown are schematic, not real temple data.

Academic — MSIT minor project

Team of three · second review, Sep 2026

E-05AcademicMinor projectMSIT

Low-latency voice-to-voice agent

Start thinking before the user finishes. Stop talking the moment they interrupt.

In progress · architecture complete, integration next, latency evaluation upcoming

With Saksham Garg and Krishan Kishore · Mentor: Mr. Akshay Singh

L1

A voice agent designed around overlap: each stage starts before the previous one has finished, and playback can be cancelled mid-sentence.

L2

Audio arrives in 20ms frames. Silero VAD detects speech. Streaming STT emits partial transcripts. When a partial is stable, the LLM starts a speculative prefill. Tokens are cut into clause-level chunks and streamed to TTS, so playback begins on the first clause.

If the user speaks, playback is cancelled and history is truncated (barge-in).

Audio frames
20ms
Per-stage tracing (evaluation pending)
p50·p95·p99
Planned API spend
₹0
  • LiveKit
  • FastAPI
  • asyncio
  • Silero VAD
  • Deepgram Nova-3
  • Groq Llama
  • Deepgram Aura-2
Illustrative timing, not to scale
USER AUDIOSILERO VADSTT PARTIALSLLMSPECULATIVE PREFILLTTS CLAUSESPLAYBACKLAST USER FRAME → FIRST RESPONSE FRAME

Prefill starts on a stable partial, before the user finishes. If the partial changes, it is discarded and reissued.

World 03 · Investigation and experiment

Research

Experiments, written up and published. Co-authored with faculty at MSIT.

EXP-001Springer LNNS 2052ICICC 2026pp. 437–444

Self-healing multilingual NLP-to-SQL

When the SQL fails, the system remembers why.

Prabhjot Kaur, Krish Arya, Kashika Malhotra, Vibhuti Mehta, Khushi Sehrawat, Preeti Singh

doi:10.1007/978-3-032-30909-9_34 ↗
L1

Ask a database a question in Hindi, Tamil, Bengali, Urdu or English and get SQL back. When the SQL fails, the failure becomes memory.

L2

A LangChain agent on Gemini 2.0 Flash translates the question, retrieves relevant schema chunks from ChromaDB, writes SQL and executes it through SQLAlchemy.

On failure, the failed query, its schema and the exact error are embedded together and stored. Later failures retrieve the most similar past cases and use them to correct the SQL.

Accuracy on Spider
92%
Execution failures auto-corrected
78%
Est. compute saved at scale
20–25%
  • Gemini 2.0 Flash
  • LangChain
  • ChromaDB
  • SQLAlchemy
  • Embeddings

--/09Failures are stored like documents and retrieved when they happen again.

Payload appears here as the loop runs

Example question, schema and SQL are illustrative. Metrics are from the published paper.

EXP-002Intl. J. of System Assurance Engineering & ManagementSpringer

Proactive financial scam detection

Flag the scam while it's still small.

Prabhjot Kaur, Vibhuti Mehta, Krish Arya

L1

A multimodal framework asking whether AI can detect a financial scam before it goes viral.

L2

Signals come from four families (text, metadata, network structure and temporal dynamics) across phishing emails, blockchain Ponzi data and social-media scam posts.

Models range from Random Forests and LSTMs to graph neural networks and LLMs. They're evaluated on classification and on early warning: lead time to virality and the cost of false alarms.

Shares at detection
~50
Viral threshold
500+
F1 (best hybrid)
0.97
  • LLMs
  • Chain-of-thought
  • Temporal GNNs
  • LSTMs
  • Random Forests

Signalstext · metadata · network · temporal

ModelsRF · LSTM · GNN · LLM

BestLLM + CoT + temporal GNN

1101001000TIME →SHARESVIRAL THRESHOLD · 500+FLAGGED · ~50 SHARESLEAD TIME

Curve is illustrative. The ~50 and 500+ thresholds, F1 0.97 and early-warning recall 0.94 are reported in the paper.

EXP-003Book chapter 8Explainable AI in Critical DomainsNova Science Publishers

XAI in autonomous systems and robotics

A robot's explanation is only useful if it arrives in time to act on.

Prabhjot Kaur, Vibhuti Mehta, Krish Arya

“Explainable AI for Safety-Critical Decision Making in Autonomous Robotics: From Black Boxes to Transparent Agents”

L1

A book chapter on explainability for autonomous systems that make safety-critical decisions.

L2

It examines XAI across autonomous vehicles, medical robotics and industrial automation, and proposes a framework for multi-modal explanation generation under robotic constraints: limited compute, real-time operation, and different explanation needs for different stakeholders.

Domains: vehicles, medical, industrial
3
Nova Science
Ch. 8
  • Explainable AI
  • Robotics
  • Human-robot interaction
  • Safety

--/04Post-hoc explanation arrives too late for a robot. The chapter argues for interpretability by design.

EXP-000Earlier explorationML / DLTime series

Environmental time-series modelling

Earlier work. Where the research started.

L1

Earlier predictive research on environmental and climate-related data.

L2

Classical and deep time-series models (SARIMA and BiLSTM), with Optuna for hyperparameter search, on environmental datasets.

  • SARIMA
  • BiLSTM
  • Optuna
  • Time series

--/05Where the research started.

World 04 · AI inside a real engineering org

Integrate

Bringing AI into how an engineering team already works, where adoption depends on trust more than on capability.

I-01EaseMyTripMiraPilot: one repository

AI-assisted code review pilot

AI inside the engineering loop, not instead of the engineer.

Pilot and workflow integration. Mira is a third-party product; I did not build it.

L1

Part of piloting AI-assisted code review at EaseMyTrip with Mira, starting from a single repository, and integrating it into the engineering workflow.

L2

The AI does a first pass and flags obvious issues, edge cases and inconsistencies. An engineer evaluates each finding and accepts or ignores it, which leaves human reviewers more time for the decisions that need them.

Repository to start
1
Final decision on every finding
Human
  • Mira
  • Code review
  • Engineering workflow
  1. 01Codebase
  2. 02Mira
  3. 03Findings
  4. 04Engineer
  5. 05Accept / ignore
  6. 06Workflow
  • edge casePossible None returned from get_fare() is dereferenced
  • nitTrailing whitespace on line 212
  • reliabilityHTTP call to pricing service has no timeout
  • stylePrefer f-string over .format()
  • inconsistencybookingId and booking_id used for the same field
  • nitConsider renaming variable `x`

Signal

—

0 accepted / 0 reviewed

Triage the findings the way an engineer would.

Simulation with example findings, to show why review quality beats issue count. Not pilot data.

05 · System log

Log

  1. 2026 · Jun 2026 – present

    EaseMyTrip

    Technology Trainee (AI/ML) · New Delhi

  2. 2026 · Apr – Jun 2026

    Tov AI

    AI Engineer · Remote

  3. 2025 · Sep – Dec 2025

    Growthic

    AI Generalist Intern · New Delhi

    • AI agentsTone-adaptiveContent agents that adapt to tone, intent and writing style for LinkedIn posts
    • AnalyticsAWS EC2Reddit campaign management and LinkedIn analytics platforms
  4. 2025 · May – Aug 2025

    Growify LLP

    AI & Automation Intern · New Delhi

    • Automation−75%Image-organization time (line-sheet mapper, batch resizer, image-similarity search)
    • Ops effort−45%Manual operational effort
    • RAGOnboardingScattered onboarding documents made queryable

06 · Topology

Stack

No proficiency bars. Each skill links to the systems where it was actually used. Hover either side.

AI systems

  • ├──LLMsB-01 B-02 B-03 E-01 E-02 E-04 E-05 EXP-001 EXP-002
  • ├──Agents & agentic workflowsE-04 EXP-001
  • ├──RAGE-02 B-01 EXP-001
  • ├──Tool callingE-04 EXP-001
  • ├──Multi-model orchestrationB-01 E-01 E-05
  • ├──EvaluationE-03 B-01 E-01 B-02 EXP-002
  • └──LangChain · LangGraph · CrewAIEXP-001

Retrieval

  • ├──pgvectorE-02
  • ├──ChromaDBEXP-001
  • ├──FAISSresume
  • ├──EmbeddingsE-02 B-01 EXP-001
  • ├──ChunkingEXP-001
  • ├──Rerankingresume
  • ├──Knowledge-base designE-02
  • └──Live search · SPARQL · scrapersE-04

Speech

  • ├──Streaming STT / TTSE-01 E-05
  • ├──LiveKitE-01 E-05
  • └──VAD · barge-inE-05

Backend

  • ├──PythonB-02 E-01 E-02 E-03 E-04 E-05 EXP-001
  • ├──FastAPIE-03 E-05
  • ├──FlaskB-03
  • ├──Async I/O · worker pipelinesB-02 E-05
  • ├──Postgres · SupabaseB-01 E-02
  • └──SQL · SQLAlchemyEXP-001

Models

  • ├──PyTorch · TensorFlowEXP-002
  • ├──Scikit-learnEXP-002
  • ├──Graph neural networksEXP-002
  • └──SARIMA · BiLSTM · OptunaEXP-000

Infra & front end

  • ├──DockerE-02 EXP-002
  • ├──n8n · OpenRouterB-01
  • ├──AWS · GCPresume
  • └──React · Next.jsE-03 B-03

“resume” = listed on the resume, not shown in a system on this site.

07 · Operator

Whoami

krish@system:~$ whoami --verboseyaml

1# operator.yaml
2name: Krish Arya
3role: AI Engineer · Builder · Researcher
4now: Technology Trainee (AI/ML) @ EaseMyTrip
5location: New Delhi, IN # UTC+05:30
6education: B.Tech IT, MSIT · 2023–27 · cgpa 8.5
7pattern: "generate → verify → feed back"
8stack: ["python", "fastapi", "pgvector", "langchain", "livekit"]
9off_hours:
10 - president: TARK Literary Society
11 - cultural_head: Prakriti MSIT # ENVA Fest · 10,000+ footfall · 500+ team
12accepting: ["freelance", "collaborations", "roles"]
13

krish@system:~$ ps --user krish8 procs

PIDSTATPROCESS
E-01prodvoice-agent<200ms end to end
E-02prodsupport-ragself-updating KB
E-03prodeval-platform50K+ conversations
E-04deployingeasydarshangoing live
I-01buildingai-code-reviewMira pilot, 1 repo
E-05buildingv2v-agentteam of 3, integration
B-02shippedmutagen-aipip install mutagen-ai
B-01shippedbylin.appsole engineer

prod = running at EaseMyTrip · click a process to open it

08 · Contact

Trans­mission

Open to freelance builds, collaborations and roles: AI systems, RAG, agents, voice, evaluation. Tell me what you're working on.

What's this about

Opens your mail client, addressed to me.

type `help`, or try `show mutagen`
Commands: help, projects, show followed by a project name, experience, research, stack, contact.