Data scientist building GenAI agents and ML systems.

I turn messy data and ambiguous product problems into reliable pipelines, evaluated models, agent workflows, and decision-ready software.

About

Research rigor, product instincts, and production discipline.

I am a Drexel MS Data Science graduate with a computer engineering background and experience spanning AI research, applied analytics, agricultural forecasting, GenAI agent systems, and consumer software platforms.

Recent work includes AgentOps CRM tooling, CommonCrawl preprocessing for LLaMA-inspired corpus preparation, XGBoost forecasting on GCP, Reddit mental-health NLP, and hybrid recommendation systems.

01 0% server response-time reduction at gurado
02 0% crop spoilage reduction with XGBoost forecasting
03 0% accuracy on Reddit mental-health classification
04 0+ students supported through Python, SQL, and Power BI labs
Featured Projects

Agent systems, ML pipelines, and applied analytics.

GitHub profile
GenAI CRM Copilot FastAPI + React + LangGraph

AgentDesk Enterprise AgentOps

Built a full-stack CRM copilot for sales teams that routes natural-language requests through a LangGraph workflow, calls safe MCP CRM tools, returns grounded answers, and tracks production-style AgentOps metrics.

  • Exposes typed CRM tools through the official MCP Python SDK over Streamable HTTP.
  • Tracks latency, p95, token usage, estimated cost, tool success, eval pass rate, and groundedness.
  • Ships React/TypeScript chat, tool traces, observability, CRM import, and Cloud Run deployment flow.
PythonFastAPIReactTypeScriptLangGraphMCPSQLiteOpenAI
NLP / ML 28K+ posts

NeuroNexus - Reddit Mental Health NLP

Built a Reddit NLP pipeline with PRAW, spaCy, NLTK, TF-IDF, and scikit-learn to classify anxiety, depression, and ADHD posts while estimating severity from weak-supervision labels.

  • 0.827 weighted F1 and 82.87% accuracy with XGBoost.
  • Severity model reached 91.62% within +-1 severity point.
Discuss this project
Recommenders 250K+ ratings

Hybrid Recommendation System

Benchmarked ALS, SVD, PCA, k-NN, and Jaccard similarity on MovieLens data to handle sparse interactions and cold-start recommendation scenarios.

  • Validated models across 59K movies.
  • Achieved RMSE between 0.668 and 0.912 across runs.
View related work
Forecasting Finance

Stock Price Prediction using Machine Learning

Modeled 10 years of market data for Apple, Boeing, IBM, Intel, and Microsoft using linear regression, decision trees, and random forests.

  • Random Forest produced the lowest MAE, RMSE, and SMAPE.
  • Used MinMaxScaler and missing-value handling for stability.
Project context
AI Prototype Agriculture

CropGuardAI

Developed an AI-driven crop-risk tool with weather analytics, predictive logic, and a lightweight Flask application for real-time localized farm insights.

  • Delivered a functional prototype under hackathon constraints.
  • Combined Python, Flask, Pandas, ML, HTML, CSS, JavaScript, and Git.
Ask about it
Analytics Tableau

Understanding India's Digital Commerce via UPI Data

Analyzed monthly UPI transaction data across Indian states to uncover digital payment adoption, transaction volume, user-count, and frequency patterns.

  • Built Tableau views for regional and temporal adoption patterns.
  • Produced insights for business and policy stakeholders.
Explore GitHub
Experience

Where I have delivered impact.

Jun 2025 - Present

Research Assistant

Drexel University College of Computing and Informatics

Built research-grade preprocessing infrastructure for LLaMA-inspired corpus preparation, with emphasis on data quality, repeatable cleaning stages, and transparent diagnostics.

Built

  • Python pipeline for CommonCrawl WARC.WET ingestion, cleaning, deduplication, and JSONL output.
  • Streaming reports that made corpus quality checks easier to inspect and reproduce.

Approach

  • Used staged filters for language detection, text quality, duplicate detection, and readability signals.
  • Added diagnostics with langid, NLTK, TextBlob, textstat, and MinHash.

Impact

  • Processed 30K+ raw records into 10K+ quality-filtered documents.
  • Improved research review by exposing intermediate counts, rejection reasons, and corpus statistics.
Jun 2024 - Jun 2025

Teaching Assistant

Drexel University College of Computing and Informatics

Supported applied data science coursework by helping students turn code, queries, and dashboards into technically correct and clearly communicated analysis.

Led

  • Python, SQL, and Power BI labs for 60+ students across applied data science coursework.
  • Hands-on debugging sessions covering data cleaning, query logic, and visualization choices.

Reviewed

  • Code submissions, SQL queries, dashboards, and written analysis for clarity and correctness.
  • Student workflows for reproducibility, interpretation, and stakeholder-ready presentation.

Outcome

  • Helped students strengthen debugging habits across Python, SQL, and BI tooling.
  • Reinforced practical data storytelling through structured feedback on dashboards and writeups.
Jun 2024 - Dec 2024

Data Scientist

SmartAgri AgroVille - Smart Agricultural Solutions

Developed forecasting and reporting workflows for crop-risk decision support, connecting model output with operational reporting for agricultural stakeholders.

Built

  • XGBoost forecasting pipeline deployed on GCP for crop spoilage risk analysis.
  • SQL and Python feature pipelines for weather, crop, and operational data inputs.

Improved

  • Lowered RMSE from 0.52 to 0.43 through feature engineering and model iteration.
  • Reduced crop spoilage from about 18% to 12.5% through improved forecasting support.

Delivered

  • Automated Power BI reporting for recurring stakeholder visibility.
  • Reduced manual reporting effort from about 12 hours/week to 3 hours/week.
Sep 2021 - Aug 2023

Software Engineer

gurado GmbH - Digital Gift Card Platform

Improved platform performance and frontend delivery for a digital gift card product serving a large active user base.

Optimized

  • PHP APIs, SQL queries, caching layers, and CDN usage for 100K+ active users.
  • Backend and page-load paths that affected high-traffic product workflows.

Modernized

  • Led Angular-to-React migration work across product-facing interface areas.
  • Improved code review and release coordination during frontend modernization.

Impact

  • Reduced server response time from about 900 ms to 270 ms.
  • Reduced page load time from 6.5s to 1.1s and improved releases from biweekly to weekly.
Skills

Tools I use to move from raw data to shipped decisions.

Languages & ML

PythonSQLPySparkRScikit-learn XGBoostRandom ForestNLPTF-IDFspaCy

Agent & App Systems

FastAPIReactTypeScriptLangGraphMCP OpenAI APIFlaskAPIsGitHubCode Review

Data, BI & Cloud

Power BITableauLookerBigQueryDatabricks AWSGCPETLData PipelinesValidation

Methods

ForecastingModel EvaluationA/B TestingHypothesis Testing RegressionClassificationRecommendersEDA
Education

Drexel-trained data scientist with a computer engineering base.

Master of Science in Data Science

Drexel University College of Computing and Informatics / GPA 3.76 / 2023 - 2025

Relevant work: computer vision, data acquisition, information visualization, applied cloud computing, AI, ML, and software project management.

Bachelor of Engineering in Computer Engineering

Savitribai Phule Pune University / 2017 - 2022

Leadership & Recognition

Communications Chair, Drexel Graduate Student Association. Member, Upsilon Pi Epsilon honor society.

Certifications

Recent learning aligned with cloud data science and applied ML.

OCI 2025 Certified Data Science Professional OCI 2025 Certified Generative AI Professional SQL for Data Science - UC Davis, Coursera Exploring Data Transformation with Google Cloud Fraud Detection on Financial Transactions with ML on Google Cloud
Contact

Open to data scientist, machine learning, analytics, and software engineering roles.

I am based in Philadelphia and open to relocation. Recruiters and hiring managers can reach me directly by email, phone, LinkedIn, or GitHub.