AWS-certified Data Scientist & ML Engineer applying statistics, experimentation, and MLOps to turn messy data into decisions across banking, healthcare, and airline products. I design and deploy fraud detection, forecasting, and analytics solutions that reduce risk, unlock growth opportunities, and give stakeholders clear, actionable insights. I enjoy working end-to-end, from exploring raw data and shaping features to shipping production-grade ML pipelines and dashboards that teams actually use.
print(RakeshSarmaKarra.masters) | Data Science & Advanced Analytics
print(RakeshSarmaKarra.university_name) | University of North Texas
print(RakeshSarmaKarra.focus) | AIML Engineer, Data Science, Data Analytics, MLOps, Cloud Analytics print(RakeshSarmaKarra.domains_worked) | Banking/Finance, Airlines, Healthcare
print(RakeshSarmaKarra.location) | United States
Portfolio - Data & AI Projects: https://rakeshsarmakarra.github.io/
Linear & logistic regression, regularization (Ridge, Lasso)
Tree-based models (Decision Trees, Random Forest, XGBoost, LightGBM)
Clustering (K-Means), KNN, SVM, Naive Bayes, basic neural networks
Time series (ARIMA, SARIMA), anomaly detection, A/B testing
Feature engineering, cross validation, hyperparameter tuning (grid/random search), SHAP
Exploratory data analysis, hypothesis testing, KPI design
Root cause analysis, recommendations, competitor analysis
Prescriptive analytics, business improvements
Agile/Scrum, user stories, project planning, milestones
Work breakdown structure, risk register, risk mitigation, communication plans
30–70% rule, quality management
Analytical decision making, problem solving
Data storytelling, written & verbal communication, team collaboration
- Developed a GenAI-powered Financial Advisor Knowledge Assistant using OpenAI GPT-4o, Claude, LangChain, and Pinecone to provide contextual responses from investment and retirement planning documents.
- Built Retrieval-Augmented Generation (RAG) pipelines leveraging embedding models and vector search, improving information retrieval relevance and reducing manual research effort for financial advisors.
- Built end-to-end Python/SQL pipelines to model program engagement and outcomes, improving predictive performance (F1, AUC) and reducing data processing time by ~30% through optimized workflows versioned in GitHub.
- Contributing as an ML Engineer volunteer to design and prototype machine learning models in Python for Murphy Charitable Foundation’s new application supporting vulnerable communities (e.g., child sponsorship and donor engagement use cases).
- Building and iterating on data pipelines using Python and SQL (data cleaning, feature engineering, basic model training), with experiments and notebooks version controlled through Git and regularly pushed to GitHub.
- Conducted AI bias research across ChatGPT, Gemini, Claude, and Meta AI, running comparative bias/hate-speech/sentiment experiments and user-trust surveys (~68% positive) while mentoring 120+ students on GCP-based ML pipelines using Python, BigQuery, and Vertex AI.
- Find my research work in below sections
- Developed an XGBoost engagement model (ROC-AUC 0.76) with SHAP-based interpretability to identify low-engagement Medicare Advantage segments, projecting up to 40% outreach uplift and translating findings into a STAR-ratings-linked business storyboard.
- Built automated Python/Snowflake pipelines and ARIMA/Random Forest forecasting models to predict baggage volumes and congestion, cutting reporting time by ~60% and delivering Tableau-based insights presented to HQ stakeholders for capacity planning.
- Automated legacy AML/fraud ETL pipelines in SAS and Python during a CMR-to-MDM migration, cutting file processing time by 40% and improving data stewardship reliability by 76%, earning Silver and Bronze recognition for regulatory-compliant delivery.
- Built and deployed predictive risk and customer-segmentation models (Scikit-learn, SAS Enterprise Miner, K-Means) across 4M+ records, improving underwriting accuracy by 12%, lifting cross-sell conversion by 22%, and cutting reporting time by 60% through automated SAS/SQL pipelines.
Click on the image to view the project
Selected professional trainings and industry-recognized learning programs focused on Generative AI, agentic workflows, responsible AI, and machine learning.
Learning Repo | Capstone Project | Certificate Project Repo | LinkedIn Post | Credly Badge
Credly Badge LinkedIn Post | Credly Badge
- Received Silver Award in Citi bank for California Consumer Privacy Act 2020, CMR project.
- Received Bronze Award in Citi Bank for CMR to MDM migration project.
- Received Work Excellence Award in ICICI bank
- Completed the Texas Higher Education Coordinating Board's AI Professional Development Program (Sept–Dec 2024), gaining hands-on competency in prompt engineering, custom GPT development, AI-enhanced content creation, ethical AI considerations, and trustworthy generative AI frameworks, aligning technical skills with responsible AI deployment strategies for academic and industry applications.
- Collaborated with a cross-functional team of four graduate students to investigate bias propagation, hallucination patterns, and fairness concerns in LLM-based systems, conducting systematic prompt-response experiments, behavioral analysis, and statistical evaluation to quantify model variability and inform design recommendations for more transparent and equitable generative AI applications in educational and research contexts.
- Participated in speaker sessions and case-based workshops on business analytics, data visualization, and predictive modeling, gaining exposure to real-world applications and tools.
- Collaborated with peers in analytics challenges and networking events, strengthening problem-solving, presentation skills, and industry connections.
- Attended the UNT Data Science Talk Series featuring leading experts, rising stars, and academic leaders presenting on cutting-edge topics including machine learning, artificial intelligence, data visualization, big data analytics, and ethical considerations in data usage, broadening exposure to industry best practices and emerging research trends.
- Participated in collaborative learning sessions that fostered knowledge exchange on state-of-the-art methodologies such as deep learning frameworks (TensorFlow, Keras), cloud-based ML tools, natural language processing, and computational data science techniques, strengthening technical acumen and staying current with the evolving data science landscape.
- Participated in workshops on Hugging Face models, including practical tokenization sessions emphasizing LLMs for real-world AI applications.
- Promotes collaborative projects, research, and learning in AI/ML open to all majors, building skills in model deployment and innovation.
- Engages members in events like online workshops that explore AI usability, such as fine-tuning LLMs for tasks like natural language processing and ethical AI use.
- Awarded a Participation Certificate for the online workshop “How to Use Hugging Face AI Models?” in Feb 2026, recognizing active engagement in LLM and AI usability sessions (certificate: Link).



