DATA × AI × ANALYTICS × SOFTWARE
Mandeep Kumar Roshan · B.Tech CSE · Data Science / AI-ML
Most of what I build starts with a dataset I don't understand yet. I clean it, argue with it, model it, and stop when it can answer something useful.
I'm a B.Tech Computer Science student working my way through data science, machine learning and analytics — mostly by building things and seeing where they break.
What I enjoy is the part before the model: figuring out what the data actually says, which features matter, and whether the answer is worth trusting. The model is just the last step. After that I try to put the result somewhere a person can use it — a Streamlit app, a Power BI dashboard, a notebook someone else can follow.
Right now I'm spending time on time-series forecasting, recommendation systems and dashboards that hold up to questions.
| Data Science | Machine Learning | Analytics & BI | Building |
|---|---|---|---|
| Cleaning, EDA, feature engineering |
Regression, classification, clustering, forecasting |
Power BI dashboards, DAX, Power Query |
Streamlit apps, Python tools |
|
Content-based anime recommender. Encodes genres, themes, studios and demographics into features, then finds similar titles with KNN and cosine distance. Ships with a Streamlit app for search and a Power BI dashboard for the analytics side.
|
An end-to-end retail dashboard: KPIs and trends, category and region forecasts, anomaly flags, and demand segments. SARIMA, Prophet and XGBoost are compared for the forecast; Isolation Forest handles anomalies, K-Means the segmentation.
|
|
Historical NIFTY-50 trading data taken from cleaning and exploratory analysis through feature engineering — year, month, daily price change — and then used to compare Linear Regression against a Decision Tree Regressor on closing prices.
|
IBM HR analytics data — 1,470 employees, 35 attributes — used to work out who is likely to leave and why. Logistic Regression, Random Forest and Gradient Boosting are compared, and the findings are written up as retention suggestions rather than just scores.
|
|
120 years of Olympic results (1896–2016) reshaped in Power Query and built into an interactive Power BI report: participation trends, medal distribution, country performance and gender participation, with DAX measures behind the KPI cards.
|
Predicting house prices from property attributes such as area, bathrooms, parking and furnishing. Linear Regression against Random Forest, scored with MAE, RMSE and R², plus a look at which features actually move the price.
|
|
A step away from data work: a Tkinter desktop app that runs FIFO, LRU and Optimal on the same reference string, so you can watch where the page faults happen instead of taking the textbook's word for it.
|
Datasets I'm still poking at, notebooks that haven't earned a repository yet, and the occasional idea that didn't survive contact with real data.
|
| Languages | Python Jupyter Notebook |
| Data & ML | Pandas NumPy Scikit-learn SciPy Statsmodels XGBoost Prophet |
| Visualization & BI | Matplotlib Seaborn Plotly Power BI Power Query DAX |
| Building & shipping | Streamlit Tkinter Git GitHub |
a question I can't answer → find the data → spend longer cleaning it than expected
→ build something small → break it → fix it
→ ship it where someone can click on it → find a better question
I'm not pretending to be an expert at any of this. Each project teaches me one thing I got wrong in the last one — a leaky feature, a metric that flattered the model, a dashboard nobody could read. That's the whole point of keeping them public.
## `07` Connect
Open to data science and analytics internships, and to anyone who wants to talk about a messy dataset.

