Mohammadali Mousavireineh
Working at the intersection of machine learning, network security and education. Turning ideas into reproducible experiments and practical skills.
Why high accuracy is not always good news
A short guide to spotting data leakage and designing more reliable evaluations.
When information from the test set influences training, evaluation can become overly optimistic.
For example, fitting a scaler or feature selector on the entire dataset before splitting leaks information from the test set.
Split first. Fit preprocessing on the training data only, then apply the learned transformation to validation and test data.
For temporal data, respect time order. If records share a person or device, consider whether a group-aware split is needed.
Interpret accuracy alongside recall and F1. Keep the final test set separate from model-selection decisions.