I evaluate AI models for a living right now — and I'm building my way into building them.
Currently an AI Evaluator at RWS, spending my days deep in model outputs: catching where LLMs get it wrong, why they get it wrong, and what "good" actually looks like at scale. That's given me a pretty unfiltered view of how these systems fail in the real world — which is exactly what I'm now putting into practice on the other side of the table, building agentic AI and ML systems myself.
I'm looking to move into an AI Engineer role, or take on freelance work building agentic AI applications, RAG pipelines, and LLM-powered tools.
Plus everything two years of evaluating models has taught me about prompt design, failure modes, and what separates a demo from something you'd actually trust in production.
Agentic AI Recruitment Copilot A multi-step recruitment assistant orchestrated with LangGraph — automates resume screening and candidate evaluation across several stages instead of a single prompt-in, answer-out flow.
RAG-based QnA Feed it any PDF, get answers grounded only in that document — no hallucinated context, no reaching outside the source. Built to actually test the limits of retrieval-augmented generation, not just demo it.
LangChain AI Agent A conversational agent built with LangChain, Gemini, and Streamlit — my hands-on deep dive into how LLM agents actually work, from a single chain up to a full agent loop.
Medical Assistance Chatbot AI healthcare assistant fine-tuning BioBART-v2 with QLoRA for domain-specific conversational NLP, served via FastAPI + Streamlit.
Sharpening the gap between "I can build an agent" and "I can build an agent someone would actually pay for." If you're hiring for AI/ML engineering, or have a project that needs an agentic system or RAG pipeline built properly — let's talk.
- LinkedIn: sakshi-maurya
- Email: sakshi3maurya@gmail.com
- CV: sakshi-maurya-cv
- Based in Mumbai · open to remote work