Telecom Churn Prediction
Production-grade telecom churn prediction system: a config-driven scikit-learn pipeline with SHAP explainability, MLflow experiment tracking and FastAPI serving, trained on…
Data Engineer · AI Engineer · Data Scientist
I’m Muhammad Farooq (also known as Muhammad Farooq Shafi), a data engineer and machine learning practitioner focused on building systems that hold up under real scrutiny — not just notebooks that ran once. My work spans the full data lifecycle: ETL pipelines with automated data-quality gates, dbt-modeled analytics warehouses backed by dozens of automated tests, and ML systems evaluated honestly against real baselines rather than a single flattering metric. I’ve built and deployed projects across intrusion detection, recommendation systems, NLP triage, demand forecasting, and computer vision — each one served through a real API or interface, versioned, and tested, not left as a one-off script. I write regularly about the engineering discipline behind trustworthy data work: honest evaluation, reproducible pipelines, and deployment done right.
Production-grade telecom churn prediction system: a config-driven scikit-learn pipeline with SHAP explainability, MLflow experiment tracking and FastAPI serving, trained on…
An end-to-end marketing analytics data pipeline that ingests, transforms and models campaign data for reporting.
A knowledge-base search system built with retrieval-augmented generation for accurate, grounded question answering.
A time-series forecasting system for energy demand, built for planning and load-balancing use cases.
Graph-based analysis of organizational networks to surface influence, collaboration and communication patterns.
A classical machine learning system for detecting fraudulent transactions with a focus on precision/recall trade-offs.
A live snapshot of my GitHub activity — commits, streaks and languages.
View GitHub profile →Machine Learning Engineer (Internship) · CodeAlpha
Developed and implemented machine learning models for real-world applications using Python, TensorFlow and scikit-learn. Performed data preprocessing, feature selection and model tuning, and collaborated with cross-functional teams to integrate ML solutions into existing products.
Data Science Intern (Internship) · Arch Technologies
Assisted in data collection, cleaning and preprocessing to support machine learning projects. Developed and tested predictive models using Python and libraries including scikit-learn and pandas, and created data visualizations to communicate insights.
Data Science Intern (Internship) · Oasis Infobyte
Supported data preprocessing, feature engineering and exploratory data analysis for ongoing machine learning projects, and assisted in building and evaluating predictive models with scikit-learn and pandas.
Data Science Job Simulation (Forage) · British Airways
Scraped and analysed customer review data to uncover findings, and built a predictive model to understand the factors that influence buying behaviour.
Data Science Job Simulation (Forage) · BCG X
Completed a customer churn analysis simulation, conducting data analysis in Python (pandas, NumPy), engineering and optimizing a random forest model that achieved a 50% recall rate in predicting churn, and delivering an executive summary of findings.
Data Science Job Simulation (Forage) · Commonwealth Bank
Built data engineering pipelines to aggregate and extract insights from datasets, applied anonymization for data-privacy compliance, and proposed data analysis approaches for well-structured, efficient databases.
Data Analytics Job Simulation (Forage) · Deloitte
Completed a job simulation involving data analysis and forensic technology, created a data dashboard using Tableau, and used Excel to classify data and draw business conclusions.
Data Analytics Job Simulation (Forage) · Quantium
Developed expertise in data preparation and customer analytics using transaction datasets, identified benchmark stores for uplift testing, and created reports for informed, evidence-based strategic decisions.
Data Science Job Simulation (Forage) · Lloyds Banking Group
Completed a Forage job simulation for Lloyds Banking Group's data science team.
Bachelor's degree, Artificial Intelligence
Jul 2023 – May 2027
Focused on the design, development and deployment of AI-driven systems. Coursework and projects covered machine learning, deep learning, computer vision, natural language processing and robotics, with hands-on expertise in Python, TensorFlow and PyTorch.
Associate's degree, Data Science
Apr 2024 – Mar 2026
Developed a strong foundation in data analysis, statistics, programming and machine learning.
Associate's Degree, Business Administration and Management
Jun 2025 – Jun 2026
Participated in business case studies, management discussions and leadership development activities focused on strategy, marketing and organizational management.
IBM · Credential ID MFXW3LL9UI6G
Google · Jun 2025
Coursera · Jun 2025
Coursera
IBM · Jun 2025 · Credential ID FVMJDQSAUDL6
HP LIFE · Feb 2025
HP LIFE · Mar 2025
Udemy
HP LIFE · Feb 2025
HP LIFE · Feb 2025
Udemy · Nov 2024
Udemy · Nov 2024
Udemy
Design and build reliable, monitored pipelines that move data from source systems into warehouses and marts, ready for analytics.
Model and deploy star-schema warehouses on Snowflake, Databricks or DuckDB with automated data-quality testing.
Train, explain and serve ML models with MLflow tracking and FastAPI/Docker deployment.
Turn raw data into decision-ready dashboards with Streamlit, Power BI or Tableau.
Naive Bayesian classifiers are built on an assumption that’s almost always technically wrong. Here’s why it works anyway, and when it’s the right tool.
Read more →
Plenty of lines can separate two classes of points. A support vector machine specifically finds the one with the widest possible margin. Here’s how and why.
Read more →
Gradient boosting gets described often and explained rarely. Here’s how it actually works, step by step, and the hyperparameters that genuinely matter.
Read more →I work on data pipelines, warehouse modeling, ETL/ELT, and applied ML/AI projects — from a single dashboard to a full production pipeline.
Yes — update this answer with your current availability.
Python, SQL, dbt, Airflow, Docker, and cloud platforms including AWS, Azure and GCP, alongside scikit-learn/TensorFlow/PyTorch for ML work.
Send a message and I'll get back to you soon — or reach out directly: