The problem
Businesses came to DotPy with a specific question they couldn’t answer from a spreadsheet: what will this cost next quarter, which customers are about to leave, what’s actually in this photo. Each question needed its own data, its own comparison of candidate models, and its own way of proving the answer would hold up once someone else relied on it instead of me.
My role
I worked as the end-to-end machine learning practitioner across eleven applied projects delivered at DotPy between 2021 and 2023: framing the problem with whoever asked it, preparing and exploring the data, comparing candidate models, validating on held-out data, and packaging the result — as a notebook, an API, or a Docker container — with documentation for whoever picked it up afterward.
What I built
The eleven projects covered forecasting, house-price and energy prediction, customer segmentation, a recommendation baseline, sentiment and emotion classification, a rule-supported chatbot, feed-forward networks on structured data, fraud detection under class imbalance, and CNN-based image classification. The subject changed every time; the delivery discipline underneath it didn’t.
Profit forecasting for startups
- Problem
- Early-stage businesses had limited, noisy historical data and needed a forecast that communicated uncertainty instead of a false-precise number.
- Approach
- I framed the forecast horizon, engineered time and business features, and compared regression and time-aware candidates on held-out data.
- Built
- A repeatable preprocessing pipeline, a comparison of candidate models with error analysis, and a notebook-style forecast artifact with handover documentation.
- Lesson
- Forecast credibility comes from honest assumptions and leakage control, not a lower validation error on its own.
House price prediction
- Problem
- Housing data mixed numeric, categorical, missing, and location-sensitive fields, and the model had to stay interpretable enough to explain its own estimates.
- Approach
- I profiled the data, handled missing and categorical fields, engineered predictors, and compared regression and ensemble models on generalization.
- Built
- A reproducible tabular preprocessing pipeline and a reusable prediction notebook with model comparisons and error metrics.
- Lesson
- Location-based features and target leakage need close checking, because they can make a housing model look stronger than it generalizes.
Net energy prediction for power plants
- Problem
- A plant's output depended on correlated physical and operating variables, and a naive random split could hide unstable relationships between them.
- Approach
- I inspected distributions and correlations, cleaned invalid readings, engineered stable predictors, and compared regression models on residual behaviour.
- Built
- A telemetry preprocessing and feature-selection workflow and a reusable prediction artifact with residual analysis.
- Lesson
- Operational telemetry needs time-aware validation, since aggregate error metrics can hide drift and shifting operating conditions.
Customer segmentation analysis
- Problem
- Raw customer and transaction data didn't turn into actionable segments on its own; the groups had to be stable and explainable to be useful.
- Approach
- I built customer-level behavioural features, standardized the feature space, compared clustering configurations, and translated the result into readable profiles.
- Built
- Customer-level feature aggregation, clustering experiments with internal quality checks, and a documented segmentation output for business discussion.
- Lesson
- A mathematically tidy cluster only matters when it's stable and connects to a real business action.
Big-market recommendation system
- Problem
- A large catalogue created discovery friction, while sparse interactions and cold-start items limited a purely collaborative approach.
- Approach
- I built content and co-occurrence features from item and interaction data, generated candidate similarities, and evaluated the ranked output.
- Built
- Catalogue and interaction preprocessing, similarity-based recommendation logic, and reusable ranked-output artifacts.
- Lesson
- A recommendation baseline should be able to explain why it suggested an item, and treat cold start and popularity bias as explicit constraints, not afterthoughts.
Sentiment analysis on textual data
- Problem
- Sentiment labels were affected by class balance, negation, and noisy or ambiguous text, and a fair comparison needed one consistent split and metric set.
- Approach
- I cleaned and normalized the text, built classical vector features, trained baseline classifiers, and compared them against a transformer-style model.
- Built
- A text-cleaning and feature pipeline, a comparison between classical and transformer-style models, and error-analysis output for handover.
- Lesson
- Macro metrics and reading the actual errors matter most when a minority sentiment class is easy to overlook.
Interactive chatbot development
- Problem
- A useful chatbot needed predictable coverage of the intents it actually supported, and a safe fallback for everything else, from a small amount of training data.
- Approach
- I defined an intent taxonomy, normalized user text, trained the intent classifier, and mapped recognized intents to response and fallback logic.
- Built
- An intent taxonomy with training examples, an intent classifier, and rule-supported response routing behind a working demo interface.
- Lesson
- A clear scope and a real fallback path matter more early on than trying to cover as many intents as possible.
Artificial neural networks for business problems
- Problem
- A feed-forward network could model nonlinear relationships in structured business data, but could just as easily overfit and be hard to justify over a simpler model.
- Approach
- I scaled the structured features, defined a compact network, trained with validation monitoring, and compared it against a conventional baseline.
- Built
- A reusable tabular preprocessing pipeline, a configurable training and validation workflow, and a baseline comparison delivered alongside the model.
- Lesson
- Extra model complexity has to earn its place through a measurable validation gain over the simpler baseline, not novelty.
Fraud detection under class imbalance
- Problem
- Fraud was rare enough that a model could look accurate overall while missing most of it, so plain accuracy wasn't a usable metric.
- Approach
- I separated the data before applying any class-balancing technique, compared weighted and resampled classifiers, and reviewed precision, recall, and threshold trade-offs.
- Built
- A leakage-aware preprocessing and splitting workflow, a comparison of imbalance-handling techniques, and confusion-matrix analysis for the trade-off discussion.
- Lesson
- Fraud models should be tuned around the cost of a missed case versus a false alarm, not around headline accuracy.
Twitter emotion prediction
- Problem
- Short social-media text carried abbreviations, hashtags, and overlapping emotions, and aggressive cleaning could strip out the signal along with the noise.
- Approach
- I normalized platform-specific noise while keeping emotion-bearing cues, compared text classifiers, and reviewed confusion between closely related emotions.
- Built
- A social-text preprocessing pipeline, a multi-class emotion classifier, and per-class confusion analysis with a reusable inference demo.
- Lesson
- Cleaning rules should keep emojis, negation, and hashtags when they carry more signal than ordinary words do.
CNN image classification
- Problem
- The image set had a limited number of samples and a risk of near-duplicate leakage between splits, which meant preprocessing discipline mattered as much as the model.
- Approach
- I built class-balanced splits, normalized and augmented the images, trained a CNN and a fine-tuned pretrained backbone, and reviewed the misclassified examples.
- Built
- An image ingestion and augmentation pipeline, a trained CNN with a transfer-learning comparison, and a confusion-based error review.
- Lesson
- Careful splitting and a look at the actual misclassified images can matter more than changing the architecture.
How it works
Every engagement moved through the same delivery lifecycle, from a scoped business question to a documented handover, whether the underlying problem was a regression, a clustering task, or a CNN.
Tech stack
- Language & data
- Modeling
- Delivery
- Docker: Packaging a notebook or model for handover
Outcome
None of the eleven projects kept its dataset, its row counts, or its final score: what survived is the lifecycle, the technology choices, and the handover discipline, not the numbers a client saw at the time. That’s a real gap in my own record-keeping, and I’d rather say so than reconstruct a metric I can’t stand behind. What every engagement did produce was a reusable notebook or packaged artifact, a comparison across candidate approaches, and documentation a client’s own team could run without me.
What I learned
Delivery discipline is portable across problems; the model is not.
The pattern that held across forecasting, segmentation, recommendation, sentiment analysis, fraud detection, and image classification was never the algorithm. It was scoping the question honestly, comparing more than one way to answer it, and handing over something documented enough to outlive the engagement.