Available for Data Science & AI Roles

Solving Complex Problems with
AI & Data Science

I build production-grade Machine Learning models and scalable Data Pipelines. Proficient in Computer Vision, NLP, and backend APIs (Flask/FastAPI).

Previously at: Cedar Gate Eydean Inc.
Puspha Raj Pandeya

The Toolkit

Data Science • Data Engineering

Based in

St. Peters, MO 🇺🇸

Open to Relocation

Technical Arsenal

Education

M.S. Computer Science
SIUE May 2026
B.E. Computer Engineering
NCIT Aug 2022
4+ Years Experience

What Managers Say

Performance Feedback

"Puspha showcased exceptional skills... His expertise in Pandas and Python allowed him to create efficient data cleaning and transformation scripts. I have no hesitation in recommending him."

Prarup Acharya

CEO, Dekods Inc. Pvt. Ltd.

"During his tenure as Data Engineer, his performance was Excellent. He held positions from Associate Data Engineer to Data Engineer."

Sigma Ghimire

HR Officer, Cedar Gate Technologies

Selected Projects

AI & Engineering Case Studies
Computer Vision & AI

Brain Tumor Segmentation (U-Net)

Built a deep learning model to automate glioma segmentation in MRI scans. Benchmarked U-Net against SegNet, achieving 64.83% Dice Coefficient for improved diagnostic precision.

PyTorch CNNs OpenCV
Machine Learning & Data Engineering

E-Commerce Predictive Analytics Pipeline

Built a scalable ETL pipeline to process 109M raw events into 23M structured sessions. Conducted predictive model bake-offs (XGBoost, LightGBM) to forecast cart abandonment, achieving a champion F1-score of 0.9697.

Python XGBoost LightGBM Parquet scikit-learn
NLP & Deep Learning

Production Sentiment Analysis System

Developed a sentiment classification system for business reviews. Fine-tuned BERT Transformers achieving 92% accuracy, outperforming traditional CNN models (85%).

Transformers TensorFlow NLP
M.S. Thesis • Research

Algorithmic Bias in LLM Code

Replicating and extending the "Bias Unveiled" paper. Developed a "Smart Combination Testing" framework that solved predicate masking, increasing bias detection by ~108% (559 to 1,164 cases) compared to the original baseline.

Python Gemini API Audit Frameworks

Work Experience

My Professional Journey

Graduate Research Assistant

Southern Illinois University Edwardsville August 2024 - Present

Lead developer of a research-focused test automation framework in Python to systematically quantify discrimination and logic errors in AI-generated code. Simulated 3.5B demographic profiles using Monte Carlo sampling.

  • Lead developer of a research-focused test automation framework in Python to systematically quantify discrimination and logic errors in AI-generated code.
  • Engineered a testing engine that simulated 3.5 billion demographic profiles using Monte Carlo sampling to detect predicate masking and logical inconsistencies in automated decision pathways.
  • Generated descriptive charts and visual comparisons to analyze algorithmic failure rates and compare bias profiles between models including Gemini and Grok.
  • Documented testing protocols, system configurations, and research methodologies to support ongoing academic publications and technical reports.

Data Engineer & Analyst

Cedar Gate Technologies June 2023 - August 2024

Engineered high-throughput ETL data pipelines and automated validation frameworks for a US healthcare platform covering 3.2M lives and $91B in spend. Utilized Python, SQL, and AWS S3 to clean, normalize, and resolve demographic data discrepancies.

  • Managed the ingestion and transformation of high-volume healthcare data for a platform covering 3.2M lives and $91B in medical spend, ensuring data readiness for downstream analytics.
  • Engineered an Automated Data Validation Framework using Python and SQL to systematically audit data integrity, replacing manual schema checks and correcting over 500 monthly data quality discrepancies.
  • Refactored and optimized legacy ETL pipelines processing high-volume relational data, applying query tuning that reduced client data onboarding time by 20%.
  • Implemented statistical matching algorithms to resolve inconsistencies in member demographic data, creating unified patient records from fragmented sources to improve risk modeling.
  • Cleaned, transformed, aggregated, and mapped raw patient data to a company-wide schema to normalize, augment, and enrich data into a single data lake source of truth.
  • Maintained strict compliance with HIPAA data security standards, implementing access controls and data masking techniques to protect sensitive Patient Health Information (PHI).
  • Collaborated with cross-functional engineering teams in an Agile/Scrum environment to troubleshoot system bottlenecks, refine technical requirements, and define data specifications.

Contract Data Analytics Consultant

Dekods INC. Jan 2023 - Feb 2023

Architected automated data cleaning pipelines using Python (Pandas) to ingest and standardize 100,000+ UNICEF field survey records. Developed interactive Power BI dashboards with DAX to visualize operational metrics for 50+ remote communities.

  • Cleaned and processed 100,000 record files for a UNICEF initiative to upgrade various infrastructures in Nepal's rural communities, reducing manual data entry effort by 90%.
  • Performed exploratory data analysis using Python and Pandas, and designed automated data cleaning and transformation scripts.
  • Created and implemented a Power BI dashboard for the UNICEF project team to display important operational metrics, utilizing DAX and visualization best practices to deliver insights.
  • Transformed raw survey data into meaningful features to assess the development status of 50+ remote communities, translating complex technical data into clear visual reports.

Associate Machine Learning Engineer & Data Analyst

Eydean INC. Nov 2021 - Jan 2023

Deployed predictive sales models (ARIMA/LSTM) and recommendation engines (KNN). Architected real-time business KPIs dashboards using the ELK Stack and Metabase, developed geospatial delivery maps, and built full-stack workflows using FastAPI and Docker.

  • Developed and deployed sales forecasting models (ARIMA, SARIMA, LSTM) for 12 warehouses, reducing inventory stockouts by an estimated 18%.
  • Engineered customer service chatbots using NLP (BERT) and Neural Networks, handling 500+ weekly queries with 85% intent recognition accuracy.
  • Designed and deployed real-time sales analysis dashboards using the ELK Stack to visualize KPIs and monitor live business trends.
  • Led a database migration project transferring enterprise client legacy SQL data to NoSQL (MongoDB), developing custom scripts to denormalize and map complex structures with zero data loss.
  • Served models via custom scalable REST APIs built in FastAPI and Flask, containerizing full application suites with Docker for standardized deployment.
  • Built recommendation systems using Cosine Similarity and KNN collaborative filtering techniques to integrate intelligent features into CRM platforms.
  • Built logistics fleet tracking/delivery visualization maps utilizing Leaflet JS, Mapbox, Folium, osmnx, and Dijkstra's algorithm to calculate optimal routes.
  • Developed "Bookgara," a full-stack on-demand marketplace, managing database schema designs, booking workflows, and core backend logic in Flask.
  • Performed extensive data analysis and designed Metabase dashboards to communicate ERP client business trends to stakeholders.

Let's Talk Data

Get in Touch

Location

St. Peters, MO

Connect & Downloads