Aabha AI Academy All articles

Is Your Dataset Ready for Machine Learning

Check labels, leakage, missingness, and splits before a first machine-learning experiment.

Outcome: Decide whether a dataset is ready for a first modelling experiment and write a concrete repair plan.

Prerequisites: Familiarity with spreadsheet rows and columns. No coding required.

Before choosing an algorithm, check whether each row represents the right event, the inputs existed when a prediction would be made, and the outcome can be trusted.

Start with a prediction sentence

For our fictional support desk: “When a ticket arrives, predict whether it will remain unresolved after 48 elapsed hours, using information recorded at arrival.” State the prediction moment, outcome, unit, and information cutoff this precisely.

The grain is one row per ticket. The target, or label, is 1 for unresolved at 48 hours and 0 otherwise. Input columns are features. Several updates to one ticket must not become separate arrival examples.

Audit a synthetic export

These figures are invented for this lesson. The export covers January through June and was extracted on 1 July.

Check Finding
Exported rows 1,200
Extra duplicate copies 20
Unique tickets 1,180
Tickets with incomplete 48-hour observation windows 30
Tickets with known 48-hour outcomes 1,150
Unresolved at 48 hours 230 of 1,150
Missing arrival product category 115 of 1,150
Available fields Arrival channel, arrival category, final resolution hours, closing-agent notes

Initial decision: repair before training. The export contains duplicate rows, unfinished outcome windows, and fields that reveal the future.

Check the outcome and its timing

Confirm and remove the 20 duplicate copies. Set aside the 30 tickets whose outcomes are unknown. An open ticket aged 12 hours cannot yet be labelled “unresolved at 48 hours.” Do not replace its missing target with zero.

Define “resolved,” including how reopened tickets count, and check examples against event histories. Today's closing status may differ from the status at hour 48. Automatically generated labels still need review. Google's guide to labels explains these quality checks.

Exclude future information

Exclude final resolution hours and closing-agent notes from the inputs because they are unavailable at arrival. Resolution information can construct historical labels but must stay outside the feature set.

Verify that “arrival category” is the saved arrival value, not a later edit. Allowing information unavailable at prediction time to influence the model is data leakage. scikit-learn's leakage guidance describes why it creates misleading results.

Choose the split before learning transformations

For future incoming tickets, use earlier months for training, a later month for validation, and the latest months for the final test. Validation supports development choices; the final test stays untouched until those choices are settled. At each simulated training date, include only labels already knowable then. Allow for the 48-hour outcome window around boundaries.

If the intended use is new customers, keep each customer's tickets in one partition. A time split alone does not guarantee this. Enforce both constraints when both matter. scikit-learn documents time-based and grouped splitting.

Learn replacement values, scaling and category vocabularies from training data only, then apply those rules unchanged elsewhere.

Explain missingness and permitted use

The missing-category rate is 115 ÷ 1,150 = 10%. Investigate whether blanks cluster by channel or month. Automatically dropping these rows could remove a particular kind of customer request. Consider an explicit “unknown” category when it represents what the live system will actually receive. Check impossible timestamps and inconsistent category spellings too. Google's data-cleaning guidance covers common defects.

Confirm permitted training use with the data owner. Record the source, extraction date, access restrictions and retention requirements. Exclude unnecessary names and contact details.

Make a bounded readiness decision

After repairs, the 1,150 labelled tickets are candidates for an initial experiment, subject to split-boundary exclusions. Begin with a simple baseline: always predict the training set's most common class. DummyClassifier implements such comparisons. Also preserve the current operational rule for comparison.

Readiness does not establish that 1,150 examples are sufficient or representative. Inspect coverage of channels, products and seasons. More rows cannot repair leaked inputs or unreliable labels.

Try it

A ticket arrives at 09:00. Its channel is known immediately; an agent changes its category at 11:00. Which value belongs in an arrival-time model?

Answer: Use the channel and the category recorded at 09:00, including a genuine missing value. The 11:00 category is too late.

Further reading: Google's dataset characteristics lesson.