1.1.4 - Evaluating experiments

1.1.4 - Evaluating experiments

Evaluation is the skill of deciding how much trust to place in experimental evidence. In this lesson you will learn how to draw conclusions from data, identify anomalies and limitations, judge precision and accuracy using uncertainties, and suggest improvements that genuinely make an investigation better. This is a core OCR practical skill because written papers often ask you to justify whether evidence is good enough, not just to calculate a value.

Evidence and Conclusions

A conclusion is not just the final sentence of a practical write-up. It is a judgement about what the results show, how strongly they show it, and whether the method was good enough to answer the original question.

In OCR-style practical evaluation, a strong conclusion usually has four ingredients:

  1. It answers the question being investigated.
  2. It refers to a result, pattern, mean value, gradient, percentage difference or uncertainty.
  3. It says how confident you can be in the conclusion.
  4. It acknowledges any limitation that affects the strength of the evidence.

For example, suppose a student investigates whether the period of an oscillator changes when the amplitude is increased slightly. A weak conclusion would be: "The period stayed the same." A stronger conclusion would be: "The mean period stayed between 1.42 s and 1.44 s for the amplitudes tested, so within the resolution of the timing method there is no evidence that the small change in amplitude affected the period."

The second conclusion is better because it links the claim to data and to the resolving power of the method. It does not pretend the experiment has proved a universal law; it says what this evidence can support.

Confidence

Confidence is a qualitative judgement about how well the evidence justifies a conclusion. It depends on features such as uncertainty, repeatability, reproducibility, anomalies, validity and known limitations.

The OCR specification also expects you to understand why this matters beyond school practical work. The scientific community validates new knowledge by checking whether results are repeatable, whether other researchers can reproduce them, whether the method is transparent, whether the data support the conclusion, and whether alternative explanations have been considered. This helps maintain integrity: claims should survive scrutiny rather than depend on one unchallenged set of measurements.

Avoid using the word "reliable" on its own. It is often too vague. It is usually clearer to write:

  • the readings are repeatable if the same person using the same apparatus obtains similar results;
  • the findings are reproducible if different people or laboratories obtain similar results using equivalent methods;
  • the conclusion has greater confidence if the evidence is precise, valid and consistent with other good evidence.

Anomalies and Limitations

An anomaly is a result that does not fit the pattern of the other results and is judged not to be part of the normal variation. It is sometimes called an outlier, but that word alone is not enough: you need a reason for how you will treat it.

Anomaly

An anomalous result is a measurement in a set of results that is judged not to belong to the inherent variation of the data.

You may remove an anomalous reading before calculating a mean if there is a good reason to think it came from a procedural failure, human error or equipment problem. Examples include a stopwatch being started late, a sensor slipping, a circuit connection becoming loose, or a scale being read from the wrong mark.

You should not remove a result simply because it disagrees with a prediction or makes the graph look untidy. A surprising result might be evidence of a real effect. The careful response is to repeat that measurement, check the apparatus and record a reason for any decision to exclude the reading.

A limitation is different from an anomaly. A limitation is a weakness in the procedure or apparatus that affects the quality of the results, even if no single reading is obviously wrong.

Common limitations include:

LimitationWhy it mattersBetter evaluation language
Human reaction time when using a stopwatchAdds uncertainty to time measurements"Reaction time is a large fraction of the measured time."
Low-resolution apparatusGives a large percentage uncertainty"The smallest scale division is too large compared with the change being measured."
Zero error or calibration errorShifts all readings in the same direction"This could reduce accuracy because every reading is biased."
Parallax when reading a scaleReading depends on eye position"The scale should be viewed perpendicular to the mark."
Poor control of variablesThe investigation may not isolate the intended factor"Temperature was not controlled, so the change cannot be attributed only to the independent variable."

The best evaluation answers are specific. "There may be human error" is usually too vague. "Starting the stopwatch by hand adds a reaction-time uncertainty of about 0.2 s to each timing" is much stronger because it identifies the source and the effect.

Worked example: deciding what to do with an anomaly

A student records the time for ten oscillations of a mass on a spring:

2.84 s, 2.86 s, 2.83 s, 3.51 s, 2.85 s

The value 3.51 s is much larger than the others. If the student noted that the mass hit the bench during that run, there is a justified procedural reason to remove it before finding the mean. If there was no such reason, the student should repeat the timing at the same conditions before deciding.

Precision, Accuracy and Uncertainty

Precision and accuracy are related, but they are not the same.

Precision

Precision is the closeness of agreement between repeated measurements made under the same conditions. It is about the spread of the readings, not whether they are close to the true value.

Precision is therefore judged by looking at how close repeat readings are to one another.

Accuracy

Accuracy is the closeness of a measured value, or result derived from measurements, to the true or accepted value.

A set of readings can be precise but inaccurate. For example, a balance with a zero error might give repeated masses that are very close together but all too high. Repeating readings and taking a mean can reduce the effect of random errors, but it will not automatically remove a systematic error.

Uncertainty

Uncertainty is an estimate attached to a measurement that describes the range within which the true value is expected to lie. It is often written as a value plus/minus an absolute uncertainty, such as 12.4 cm +/- 0.1 cm.

The plus/minus part is the margin of error for that measurement or result. For 12.4 cm +/- 0.1 cm, the margin of error is 0.1 cm, so the result is being treated as lying between 12.3 cm and 12.5 cm.

OCR practical work often uses apparatus uncertainty as a sensible estimate:

  • for analogue scales, use half the smallest scale division for one reading unless the apparatus gives a different uncertainty;
  • for a distance found from two scale readings, include the uncertainty in both readings;
  • for digital instruments, use plus/minus the resolution of the display unless the question states a different uncertainty;
  • for stopwatches, the operator's reaction time may be much larger than the display resolution, so a realistic timing uncertainty may be needed;
  • if the exam gives an absolute uncertainty, use the value given in the question.

OCR also notes that uncertainty conventions in textbooks are not perfectly consistent, especially for digital instruments. In an answer, the safest approach is to state the assumption you are using and follow any uncertainty stated in the question.

Percentage Uncertainty

percentage uncertainty=absolute uncertaintymeasured value×100%\text{percentage uncertainty}=\frac{\text{absolute uncertainty}}{\text{measured value}}\times 100\%

This equation tells you how significant the uncertainty is compared with the size of the measurement. The same absolute uncertainty matters much more for a small measurement than for a large one.

Worked example: apparatus uncertainty for a length

A student measures the length of a metal rod using a ruler marked every 1 mm. The rod length is found from the difference between two readings, so the uncertainty is included twice:

  • uncertainty in each ruler reading = +/- 0.5 mm;
  • total uncertainty in the length = +/- 1.0 mm;
  • measured length = 242 mm.

Percentage uncertainty:

1.0242×100%=0.413%\frac{1.0}{242}\times 100\%=0.413\%

The length can be written as 242 mm +/- 1 mm, with a percentage uncertainty of about 0.4%.

Worked example: a small difference can have a large percentage uncertainty

A rod is measured before and after heating.

  • length before heating = 42.6 cm +/- 0.1 cm;
  • length after heating = 43.3 cm +/- 0.1 cm;
  • increase in length = 0.7 cm.

The increase is found by subtracting two readings, so the absolute uncertainties are added:

Δ(increase)=0.1 cm+0.1 cm=0.2 cm\Delta(\text{increase})=0.1\text{ cm}+0.1\text{ cm}=0.2\text{ cm}

Percentage uncertainty in the increase:

0.20.7×100%=28.6%29%\frac{0.2}{0.7}\times 100\%=28.6\%\approx 29\%

That is a large percentage uncertainty. The original length readings look precise, but the change being measured is small, so the conclusion about the increase is weak unless the method is improved.

Judging Results Against Uncertainty

Evaluation becomes stronger when you use uncertainty to decide whether a conclusion is justified. Do not just calculate a percentage and leave it there. Ask what the number means for the claim being made.

Percentage Difference

percentage difference=experimental valueaccepted valueaccepted value×100%\text{percentage difference}=\frac{|\text{experimental value}-\text{accepted value}|}{\text{accepted value}}\times 100\%

This is often what students mean by "percentage error" when comparing an experimental result with an accepted value. A percentage difference tells you how far away the result is from the accepted value. It does not, by itself, tell you why the difference occurred.

Worked example: evaluating a value for gg

A student obtains:

g=9.6±0.3 m s2g=9.6\pm 0.3\text{ m s}^{-2}

The accepted value is 9.81 m s29.81\text{ m s}^{-2}.

The student's range is:

9.3 m s2 to 9.9 m s29.3\text{ m s}^{-2}\text{ to }9.9\text{ m s}^{-2}

The accepted value, 9.81 m s^-2, lies inside this range. A good conclusion is:

"The result is consistent with the accepted value within the uncertainty. The uncertainty is fairly large, so the experiment supports the value of gg, but it is not a very precise determination."

Now compare this with:

g=8.9±0.1 m s2g=8.9\pm 0.1\text{ m s}^{-2}

The range is 8.8 m s^-2 to 9.0 m s^-2, which does not include 9.81 m s^-2. The uncertainty is small, but the result is inaccurate. This points towards a systematic error, an invalid assumption, or a limitation in the procedure.

A useful evaluation sentence has the shape:

"Because [evidence or calculation], the conclusion is [supported / not supported / only weakly supported], and the main reason is likely to be [specific limitation or uncertainty]."

Now use the same logic to judge whether a numerical result agrees with an accepted value.

Refining the Design

Refining an experimental design means suggesting changes to the procedure or apparatus that would improve the quality of the evidence. In OCR answers, the improvement should be matched to the weakness.

Weak answer:

"Use better equipment."

Strong answer:

"Use light gates instead of a hand-operated stopwatch, because this removes the operator's reaction-time uncertainty from the timing."

The strongest improvement statements have three parts:

  1. Identify the limitation.
  2. Name the specific change.
  3. Explain the effect on the evidence.

Use these common links:

Limitation identifiedSpecific improvementEffect on evidence
Reaction time is significantUse light gates or data loggingReduces timing uncertainty and improves precision
Scale reading has parallaxUse a set square/fiducial marker or read perpendicular to the scaleReduces reading error
Zero error is presentCheck and correct the zero reading, or recalibrateImproves accuracy by removing a systematic shift
Change measured is too smallIncrease the measured interval or use higher-resolution apparatusReduces percentage uncertainty
Independent variable is not isolatedControl named variables such as temperature, length or supply voltageImproves validity
Random scatter is largeTake repeats and calculate a mean after justifying any anomaly treatmentImproves precision and confidence

Be careful: "take more repeats" is useful for random variation, but it does not fix a systematic error. "Use a more precise instrument" is useful only if you name the instrument or resolution and explain what uncertainty it reduces. "Control variables" is useful only if you name the variable that matters.

Worked example: improving a timing experiment

A student times one oscillation of a pendulum with a stopwatch. The measured time is about 1.2 s, and repeated readings vary by about 0.2 s.

Two strong improvements would be:

  • Time 20 oscillations and divide by 20. The reaction-time uncertainty is then spread over a much longer time interval, reducing the percentage uncertainty in the period.
  • Use a fiducial marker at the centre of the swing and start/stop timing as the bob passes the same point in the same direction. This makes the timing point more consistent.

If suitable apparatus is available, using a light gate or motion sensor could reduce the human reaction-time limitation further. The key is that each improvement targets the actual weakness in the method.

Refining the Design Continued

That final example shows the important habit: name the weakness before naming the cure.

Good evaluation is not a list of stock phrases. It is a chain: evidence leads to a conclusion, uncertainty controls confidence, limitations explain weaknesses, and improvements target those weaknesses.