Forward Deployed AI Engineer
Axitem Software Solution Inc. · Remote, USA
- Made AI quality measurable for enterprise clients: designed and shipped continuous evaluation loops (CI/CD for LLMs) that score relevance, faithfulness, correctness and coherence of RAG and agent workflows on every change, so regressions fail the pipeline instead of reaching users.
- Turned vague "the answers are wrong" complaints into tracked, reproducible defects by building rubric-based grading pipelines in Python pairing deterministic PyTest checks with LLM-as-a-Judge, isolating and closing P1 hallucination bugs.
- Exposed retrieval failures that leaderboard benchmarks missed by curating task-specific Golden Datasets with client domain experts in backlog refinement workshops, replacing generic benchmarks with the client's own ground truth.
- Closed the loop from production back into the test suite: captured live user feedback and request traces, triaged edge cases in backlog refinement, and promoted each into the Golden Dataset so every reported failure became a permanent regression case.
- Shortened time-to-value on client AI features across 2-week sprints by catching failure modes — missed retrieval, prompt drift — before UAT rather than after release.
- Put grounded, source-cited answers in front of support teams daily: delivered RAG document-Q&A assistants (LangChain, OpenAI API) embedded inside the portals they already worked in, replacing hours of manual lookup.
- Unlocked privacy-sensitive enterprise deals by deploying evaluated open-source models (Llama 3, Mistral) entirely on client-owned infrastructure, with no client data leaving their environment.
- Cut deployment failures and reduced cross-cloud migration to a configuration change by standardizing delivery on a Dockerized FastAPI template with eval telemetry built in.