chiranjeevi@ai-eng: ~
$

Tempe, AZ  ·  Open to AI Engineering roles

Chiranjeevi Gundu

Forward Deployed AI Engineer — AI Systems Evaluation

Forward Deployed AI Engineer with an M.S. in Computer Science (AI track) and four years of professional engineering. Makes AI quality measurable for enterprise clients — continuous evaluation loops (CI/CD for LLMs) on task-specific golden datasets, graded for relevance, faithfulness, correctness and coherence, so a quality regression fails a build instead of reaching users. Ships with backend rigor: typed FastAPI services, PostgreSQL data modeling, automated tests, and secure cloud and on-prem deployment.

0%faster ticket triage
0%faster client queries
0%test coverage held
0automated tests shipped

Experience

Four years of shipping systems teams actually adopt and trust in production.

Forward Deployed AI Engineer

Axitem Software Solution Inc. · Remote, USA

Aug 2025 – Present
  • Made AI quality measurable for enterprise clients: designed and shipped continuous evaluation loops (CI/CD for LLMs) that score relevance, faithfulness, correctness and coherence of RAG and agent workflows on every change, so regressions fail the pipeline instead of reaching users.
  • Turned vague "the answers are wrong" complaints into tracked, reproducible defects by building rubric-based grading pipelines in Python pairing deterministic PyTest checks with LLM-as-a-Judge, isolating and closing P1 hallucination bugs.
  • Exposed retrieval failures that leaderboard benchmarks missed by curating task-specific Golden Datasets with client domain experts in backlog refinement workshops, replacing generic benchmarks with the client's own ground truth.
  • Closed the loop from production back into the test suite: captured live user feedback and request traces, triaged edge cases in backlog refinement, and promoted each into the Golden Dataset so every reported failure became a permanent regression case.
LangChainLangGraphPineconeFastAPIPostgreSQLLlama 3Docker

Student Worker — Administrative & Marketing Support

Sodexo · Saint Louis, MO (on-campus role during Master's)

Aug 2024 – May 2025
  • Drove AI tool adoption across the administrative team — demonstrating and training colleagues on Copilot in Office applications and ChatGPT for planning, documentation, and drafting tailored communications — so the team completed routine work faster and took on additional activities.
  • Used AI assistants (ChatGPT, Gemini, spreadsheet AI) daily to compile student dining feedback from online portals, suggestion boxes, and app reviews into themed summaries for management, draft social media content, and support inventory and stocking planning; built spreadsheet formulas and macros to speed daily entry of sales, meal-plan transaction, and temperature-log data.
  • Programmed daily digital menu boards across dining halls and retail locations, verifying dietary and allergen labeling (vegan, vegetarian, gluten-free) was accurate before every service.
AI EnablementChatGPTMicrosoft CopilotGeminiSpreadsheet Automation

Backend Software Engineer

Tata Consultancy Services (TCS) · Chennai, India

Aug 2021 – Jul 2023
  • Cut manual ticket-sorting time 40% for a Fortune 500 client over a 6-sprint timeline by integrating an ML text-classification model into their IT support system via Python (FastAPI) REST APIs.
  • Designed a confidence-threshold fallback mechanism, ensuring low-confidence predictions defaulted to the manual queue and passed strict User Acceptance Testing (UAT).
  • Improved mission-critical client workflow performance 60% by utilizing EXPLAIN/ANALYZE to profile PostgreSQL queries, adding targeted indexes, and resolving P2 timeout bugs.
FastAPIPostgreSQLScikit-learnPyTestCI/CD

Projects

Systems built end to end — data model, agent layer, API, and client.

hybrid-rag

Hybrid Retrieval Engine with a Measured Eval Harness · MIT licensed

View Repo
0.920MRR, hybrid
22golden eval cases
3interfaces
0API keys required
  • Built a hybrid document retrieval engine fusing a dense arm (BAAI/bge-base-en-v1.5, 768d, via ONNX Runtime) with a Postgres full-text arm using ts_rank_cd cover density, combined by Reciprocal Rank Fusion at k=60 over ranks rather than raw scores.
  • Proved the two arms have opposite blind spots rather than assuming it: on a committed 22-case golden set the lexical arm scores 0.000 MRR on every paraphrase query, while fusion lifts exact-identifier retrieval from 0.917 to a perfect 1.000 — 0.920 MRR hybrid against 0.898 dense and 0.273 lexical alone.
  • Measured cross-encoder reranking and shipped it disabled because it lost: −0.032 MRR overall and −0.061 on paraphrase. The code path and the measurement both stay in the repo, because a negative result you can reproduce is worth more than an untested feature.
  • Wrote a reproducible eval harness computing MRR, recall@k and nDCG@10 positionally with no LLM judge, wired to a CI regression floor so a retrieval-quality drop fails the build rather than reaching a user.

Aadyon Assist

Self-Hosted AI Life-Ops Platform · MIT licensed

View Repo
6Docker services
174automated tests
20tables under RLS
  • Architected and open-sourced a self-hosted AI platform end to end — data model, tool-calling agent layer, ingestion, REST API and installable PWA — running as six Docker Compose services on FastAPI and Postgres 16.
  • Built a bounded agent loop capped at an explicit step limit, with every tool call, its arguments and its result persisted to an append-only message log for replay and audit.
  • Engineered a human-in-the-loop safety model: read-only analysis runs autonomously, while any action with a real-world side effect becomes a proposal the user must approve, so no irreversible action executes on its own.
  • Designed multi-tenant isolation with Postgres Row-Level Security forced at the database across 20 tables, deliberately running the API as a restricted role rather than the schema owner — Postgres exempts owners from RLS, so that single choice is the difference between isolation and none.

Aadyon Worth

Zero-Knowledge Personal Finance Server · private repo

3tables the server may hold
36structural invariants
0plaintext it can read
  • Built a sync server that provably cannot read what it stores: records are encrypted client-side under a key derived from the user password with Argon2id, and the server only ever receives ciphertext.
  • Turned that guarantee into a build gate rather than a promise — a policy test walks the source and fails CI if the server imports a crypto library, if any SQL aggregates or orders by ciphertext, or if the schema grows a table beyond the three it is allowed.
  • Pinned the cross-language crypto contract in committed test vectors checked independently from Python and Swift, so a client that diverges on key wrapping or serialisation fails a test instead of silently producing records another client cannot open.
  • Designed the password-reset flow around what zero-knowledge actually costs: proving control of a mailbox cannot recover a key nobody holds, so reset destroys the records — and the page says so in those words rather than burying it.

floci-lab

Deployment Platform — Local Emulator and AWS CDK · MIT licensed

View Repo
2services on ECS Fargate
0NAT gateways
1shared RDS instance
5CloudWatch alarms
  • Wrote AWS CDK in TypeScript deploying two containerised services to ECS Fargate behind a single Application Load Balancer, sharing one RDS Postgres 16 instance with pgvector in isolated subnets that have no route to the internet.
  • Designed the network around its dominant cost: zero NAT gateways, with tasks in public subnets reachable only from the load balancer security group. A default VPC layout provisions one gateway per availability zone and would have spent more on address translation than on compute and database combined.
  • Kept every secret out of the codebase and out of CloudFormation — RDS generates its own password into Secrets Manager, the ECS agent resolves secrets by ARN at task start, and the container reaches S3 through its task role with no access key existing anywhere.
  • Collapsed two 188-line per-product deploy scripts into one parameterised script, which surfaced a bug latent in both copies: each hardcoded the same ingress-proxy name and deleted the other product’s proxy on deploy.

llmkit

Shared LLM Plumbing · Apache-2.0

View Repo
4routing tiers
6modules
819lines of library code
  • Extracted the LLM plumbing two applications had each grown their own copy of — tier-based model routing over LiteLLM, a single traced chat() chokepoint, storage and parsing helpers — so the copies stop drifting apart.
  • Kept the dependency surface deliberately thin: the core imports with zero optional extras and CI asserts exactly that, so a consumer wanting routing does not inherit boto3, litellm and pypdf along with it.
  • Inverted configuration so the library never reads a secret file — callers resolve secrets and push values in — because a secret-file convention belongs to an application, not to an LLM client.
  • Consolidating the copies fixed a real defect: one fence-stripper failed to trim whitespace before testing for the opening fence, so any model response beginning with a newline silently failed to parse and the document landed in an error state.

Synapse Storage System

Vision-Model Document Triage for a NAS Archive

View Repo
3classification methods
19automated tests
  • Built document triage over a NAS archive with a local multimodal model, preferring extracted text and using vision where a document has none.
  • Fixed a metric that was actively misleading: filename-and-extension fallbacks were recorded and counted identically to real model classifications, so a run where the model never responded looked exactly like a successful one. Classification now records which method decided — model, heuristic, or mock — and the fallback increments its own counter.

Tech Stack

What I reach for, grouped by where it lives in the system — evaluation first, since that is the job.

AI Evaluation

LLM EvaluationEval Set & Golden Dataset DesignRubric-Based GradingLLM-as-a-JudgeDeterministic Code-ChecksHallucination & Faithfulness DetectionRefusal on Low ConfidenceMRRRecall@knDCG@10Committed BaselinesRegression Gating in CIRegression HarnessesCitations & GroundingGuardrails AI

Retrieval & RAG

RAGHybrid RetrievalRank Fusion (RRF)Cross-Encoder RerankingStructure-Aware ChunkingChunking & Retrieval TuningEmbeddingsSemantic SearchVector Search (kNN)pgvectorPineconeChromaPostgres Full-Text SearchIncremental Re-IndexingDocument Parsing

Agents & Orchestration

Tool-CallingAgentic WorkflowsLangGraphLangChainLlamaIndexModel Context Protocol (MCP)Bounded Agent LoopsHuman-in-the-LoopApproval GatesAudit TrailsPrompt EngineeringContext Optimization

Models & Inference

PyTorchONNX RuntimeOn-Device & Edge InferenceOllamaLiteLLMModel Routing (Tiered)Model-Agnostic PipelinesLlama 3MistralGPT-4o / 4o-miniAnthropic ClaudeOpenAI APIOpenAI VisionHugging FaceFine-Tuning Fundamentals (LoRA, QLoRA)Scikit-learnpandasNumPyText ClassificationPrompt CachingToken & Context BudgetingData Privacy (Local Models)

Backend & Data

PythonSQLFastAPIPydanticREST APIsAsync I/O (asyncio)Node.jsReactReact NativeMicroservicesSOAPostgreSQL 16RedisRow-Level SecurityMulti-Tenant IsolationJWT AuthFernet EncryptionConnection PoolingSQL ProfilingEXPLAIN ANALYZEIndexing & Query OptimizationETL PipelinesHealth EndpointsStructured LoggingIMAPMicrosoft Graph

Cloud & DevOps

AWS ECS FargateAWS RDSAWS CDK (TypeScript)Infrastructure as CodeAWS Secrets ManagerCloudWatchIAM Identity CenterVPC DesignAWS S3ECRAzureDockerDocker ComposeDocker SecretsOn-Prem / Local DeploymentEnv-based ConfigGitHub ActionsCI/CDGitPyTestyoyo MigrationsLangFuse TracingMonitoring & LoggingTailscaleAgile/ScrumTechnical Documentation

Languages & Tooling

PythonSQLJavaScriptTypeScriptBashPowerShellClaude CodeGitHub CopilotCursor

Education

M.S. Computer Science — Artificial Intelligence Track

Saint Louis University · St. Louis, MO

Aug 2023 – May 2025

B.Tech — Electronics & Communication Engineering

Jawaharlal Nehru Technological University · India

Jul 2017 – Aug 2021

Contact

Building something with agents, retrieval, or a backend that needs to hold up? Let's talk.

Psst 👋 single neuron cell here, reporting from Chiranjeevi's brain. Poke my little memory about him 🧠