Memorizing is not learning
A student who memorizes an exam may score well without solving a new problem. Selecting robot settings from the best historical result presents a similar danger. We separate development data from evaluation data. Tables are hypothetical and photographs provide context. A profit curve alone cannot establish behavior under new conditions. The central question is whether the evaluation was truly separated from the decisions used to select the configuration, rather than merely given an impressive label.
Separate development from evaluation
Define the boundary first
The example divides twelve months into eight for development and four for evaluation. This ratio is not universally appropriate. Record the boundary and rationale before inspecting results. Define indicator warm-up and treatment of positions crossing the boundary. Repeatedly consulting evaluation data to tune settings means it is no longer an untouched final test. The label must reflect actual data use. Preserve the dates, assumptions and every revision so the testing history can be reviewed.
Example: 8 months development + 4 evaluation
Three settings and one criterion
Consider A with profit eight and drawdown nine, B with profit six and drawdown three, and C with profit five and drawdown two. The prewritten rule rejects drawdown above four, then chooses the highest remaining profit. Reject A and select B. Changing the rule after viewing the table to favor A creates a different experiment. The report should include the criterion and rejected outcomes, not just the number that makes the selected version look attractive.
A: profit8, drawdown9 B: profit6, drawdown3 C: profit5, drawdown2 → B
Freeze after selection
Evaluate the chosen version with its saved settings. If results disappoint, do not quietly replace it with whichever alternative looks better on that same segment and publish only that outcome. Further research is possible, but acknowledge that this data informed selection. Design the next evaluation separately. An unfavorable result is still decision evidence. Deleting it improves the appearance of the report, not the reliability of the program. The trial history is part of the evidence.
Freeze the version Do not hide failed evaluation
Historical forward and future demo
Forward testing in the MetaTrader tester can mean a later historical segment. A demo-forward run collects new observations from now onward. These are different and answer different questions. Document data, costs and order behavior. Out-of-sample success does not validate reconnection controls or guarantee future returns. Validation requires multiple kinds of evidence. If you compare several forward results and select the best, disclose that additional selection rather than presenting it as an untouched final evaluation.
Historical forward ≠ future demo observations
Record how many choices you made
Hundreds of trials give historical noise more chances to look impressive. Record trial counts, manual changes and criteria. A neighborhood of settings with similar behavior may be more informative than one exceptional point, but is not proof. Controlled changes and ablation tests reveal which decisions drive results. The goal is to reduce unknowns, not manufacture a guarantee. Fewer parameters alone are not enough either; selection procedure and data use remain important.
Trial count + manual revisions One exceptional point invites questions
Exercise and answer
Solve the table with maximum drawdown four followed by highest profit: the answer is B. Add an evaluation column and retain negative results. The worksheet records boundary, criterion, version and data reuse. Explain whether a failure calls for further evidence, a specific revision or stopping. Next we examine demo safety, duplicate orders and lost communication. This assignment teaches a review process; it does not predict that any selected configuration will make money.
Selection answer: B Boundary + criterion + version + history