BAMAUP · AUTO-TRADING · FREE LESSON 5 / 7

Overfitting and Out-of-Sample Tests: Keep the Exam Unseen

Actual video duration: 03:46

Video chapters

Numerical examples are hypothetical. Licensed photographs provide context; this is not live MetaTrader screen capture or BamaUp trading performance.

Memorizing is not learning

A student who memorizes an exam may score well without solving a new problem. Selecting robot settings from the best historical result presents a similar danger. We separate development data from evaluation data. Tables are hypothetical and photographs provide context. A profit curve alone cannot establish behavior under new conditions. The central question is whether the evaluation was truly separated from the decisions used to select the configuration, rather than merely given an impressive label.

Separate development from evaluation

Define the boundary first

The example divides twelve months into eight for development and four for evaluation. This ratio is not universally appropriate. Record the boundary and rationale before inspecting results. Define indicator warm-up and treatment of positions crossing the boundary. Repeatedly consulting evaluation data to tune settings means it is no longer an untouched final test. The label must reflect actual data use. Preserve the dates, assumptions and every revision so the testing history can be reviewed.

Example: 8 months development + 4 evaluation

Three settings and one criterion

Consider A with profit eight and drawdown nine, B with profit six and drawdown three, and C with profit five and drawdown two. The prewritten rule rejects drawdown above four, then chooses the highest remaining profit. Reject A and select B. Changing the rule after viewing the table to favor A creates a different experiment. The report should include the criterion and rejected outcomes, not just the number that makes the selected version look attractive.

A: profit8, drawdown9
B: profit6, drawdown3
C: profit5, drawdown2 → B

Freeze after selection

Evaluate the chosen version with its saved settings. If results disappoint, do not quietly replace it with whichever alternative looks better on that same segment and publish only that outcome. Further research is possible, but acknowledge that this data informed selection. Design the next evaluation separately. An unfavorable result is still decision evidence. Deleting it improves the appearance of the report, not the reliability of the program. The trial history is part of the evidence.

Freeze the version
Do not hide failed evaluation

Historical forward and future demo

Forward testing in the MetaTrader tester can mean a later historical segment. A demo-forward run collects new observations from now onward. These are different and answer different questions. Document data, costs and order behavior. Out-of-sample success does not validate reconnection controls or guarantee future returns. Validation requires multiple kinds of evidence. If you compare several forward results and select the best, disclose that additional selection rather than presenting it as an untouched final evaluation.

Historical forward ≠ future demo observations

Record how many choices you made

Hundreds of trials give historical noise more chances to look impressive. Record trial counts, manual changes and criteria. A neighborhood of settings with similar behavior may be more informative than one exceptional point, but is not proof. Controlled changes and ablation tests reveal which decisions drive results. The goal is to reduce unknowns, not manufacture a guarantee. Fewer parameters alone are not enough either; selection procedure and data use remain important.

Trial count + manual revisions
One exceptional point invites questions

Exercise and answer

Solve the table with maximum drawdown four followed by highest profit: the answer is B. Add an evaluation column and retain negative results. The worksheet records boundary, criterion, version and data reuse. Explain whether a failure calls for further evidence, a specific revision or stopping. Next we examine demo safety, duplicate orders and lost communication. This assignment teaches a review process; it does not predict that any selected configuration will make money.

Selection answer: B
Boundary + criterion + version + history

Your assignment

Write the time boundary and selection rule before running; evaluate the three configurations.

Worksheet fileExercise and answer guideDownload captions
Reveal the worked answer and review criteria

With drawdown at most 4, reject A and select B for the highest remaining profit. Evaluation data reused for tuning is no longer untouched.

Technical basis and image credits

MetaQuotes: OrderCalcProfit, OrderSend, Strategy Testing, Optimization and Testing Reports. Context photographs: Unsplash; credits accompany the owner publishing kit.

Educational content, not signals, guaranteed returns or personalized investment advice.

Support