Welcome To Ask or Share your Answers For Others

python - How does using an Imputer before calling Train Test Split cause Data Leakage?

posted Oct 7, 2021 in Technique[技术] by 深蓝 (71.8m points)

python - How does using an Imputer before calling Train Test Split cause Data Leakage?

I was reading content on Data Leakage which said:

For example, imagine you run preprocessing (like fitting an imputer for missing values) before calling train_test_split(). The end result? Your model may get good validation scores, giving you great confidence in it, but perform poorly when you deploy it to make decisions.

How does excluding validation data from imputing result in a better model? Shouldn't it result in the model performing poorly with the missing values?

question from:https://stackoverflow.com/questions/65902936/how-does-using-an-imputer-before-calling-train-test-split-cause-data-leakage

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…