Null vs Alternative Hypotheses: The Two Sentences That Make Your Project Provable

Rung 1 of 10 in our hypothesis testing series

A process improvement pilot can produce a lower average cycle time. But is the change real, or could the difference have occurred through ordinary process variation? In Lean Six Sigma, two carefully written statements turn that question into a testable decision.

The null hypothesis (H₀) describes the no-change position. The alternative hypothesis (H₁) states the effect or difference the project is investigating. Together, they give statistical analysis a clear purpose: assess whether the collected evidence is strong enough to reject the null.

This distinction matters because data can contradict a null hypothesis, but a test cannot prove it true. If the evidence is insufficient, the correct conclusion is “fail to reject H₀”, not “prove there was no improvement.”

What H₀ and H₁ mean in practice

Suppose a health insurer’s claims-triage process has a baseline mean cycle time of 8.6 days. The team wants to know whether a pilot has reduced the average.

For a directional, left-tailed test, the project’s claims can be framed as:

  • H₀: μ ≥ 8.6 days. The process has not achieved a reduction.
  • H₁: μ < 8.6 days. The mean cycle time has decreased.

The test calculation evaluates the null at its boundary, μ = 8.6 days. That is why H₀ is often written in shorthand as an equality to the baseline. The equality is the reference point used to calculate the test statistic; the fuller one-sided null includes values at or above that boundary.

A two-sided question, whether the mean has changed in either direction, would instead use H₀: μ = 8.6 and H₁: μ ≠ 8.6. The form of H₁ determines which results count as evidence.

Three rules for writing defensible hypotheses

  1. Anchor H₀ to the baseline. Use an equality drawn from a documented, relevant process benchmark. Be precise about the metric, unit, population, and time window.
  2. Make the pair mutually exclusive and collectively exhaustive. The statements should not overlap, and together they should cover the possible outcomes. For example, “mean below 8.6 days” and “mean at least 8.6 days” are distinct and cover all values.
  3. Choose the test direction before examining results. Decide whether H₁ is one-tailed (a specified direction) or two-tailed (any difference) before reviewing the pilot data.

The third rule protects the significance level. If a team inspects the result, then chooses whichever one-tailed direction makes it look significant, it has effectively searched both tails while using the threshold for one. At a nominal α = 0.05, that post-hoc selection can raise the false-positive rate to roughly 0.10. A one-tailed test is appropriate when the direction was set in advance and an effect in the opposite direction would not support the project claim. For context on test formulation, see Penn State’s one-sample mean t-test guide.

Worked example: claims-triage cycle time

Assume the insurer pilots a revised triage process on 32 claims. The pilot mean is 7.4 days, with a sample standard deviation of 1.6 days. The team uses a one-sample t-test against the established 8.6-day baseline, with a pre-set α = 0.05 and 31 degrees of freedom.

The standard error is:

SE = 1.6 / √32 = 0.283 days

The test statistic is:

t = (7.4 − 8.6) / 0.283 = −4.24

For a left-tailed test at this significance level, the critical value is approximately −1.697. Because −4.24 is below −1.697, the result falls in the rejection region. The one-tailed p-value is approximately 0.0001, so the team rejects H₀.

In practical terms, the pilot provides strong evidence that the mean cycle time is below the baseline, subject to the test assumptions and the quality of the data. It does not by itself establish that the intervention caused the change or that every claim will be processed faster.

Pilot test statistic compared with the critical boundary on a left-tailed t-distribution

Now consider a smaller observed improvement: the pilot mean is 8.2 days, a reduction of just 0.4 days.

The standard error remains 0.283 days:

t = (8.2 − 8.6) / 0.283 = −1.41

This result is inside the critical boundary of −1.697. The team therefore fails to reject H₀; the one-tailed p-value is about 0.084, above 0.05. This does not prove that no improvement exists. It means this pilot has not provided sufficient evidence, at the chosen threshold, to distinguish the observed reduction from noise.

That distinction can protect a business decision. If the proposed countermeasure costs $180,000 to deploy, a team should not treat a modest, statistically unconfirmed difference as proof of benefit. It can gather more data, refine the pilot, or assess operational value and risk before committing funds.

Important: This worked example treats the baseline as a fixed, relevant reference. If 8.6 days is itself an estimate from a finite sample, its uncertainty should also be accounted for, often using an appropriate two-sample comparison. The data should also meet the test’s assumptions, including independent observations and no severe outliers that distort the mean.

Type I and Type II errors: the business trade-off

A hypothesis test manages uncertainty; it does not remove it. Two possible errors frame the risk of a decision:

Error What it means Possible business cost Ways to reduce risk
Type I (false alarm) Reject a true H₀; conclude there is a change when there is not Unnecessary rollout, implementation cost, disruption, or added risk Pre-specify α and the test direction; verify measurement and assumptions; replicate important findings
Type II (missed signal) Fail to reject H₀ when a real, meaningful change exists Lost savings, continued delays, or a worthwhile improvement left unused Plan adequate sample size; reduce measurement noise; collect representative data; define a meaningful effect

For a fixed sample size and underlying effect, lowering α raises β: making a false alarm less likely generally makes a real effect harder to detect. To reduce both risks, projects often need more observations, a more reliable measurement system, or a design that reduces variation. Power is 1 − β; it is the probability of detecting an effect of the specified size when that effect is real.

Where hypotheses fit in DMAIC

In DMAIC, the hypotheses become an operating instruction in the Analyse phase. A team develops candidate causes, potential critical Xs, and then uses statistical and visual tools to test whether the evidence supports a relationship with the measured outcome, Y. The aim is to move from a plausible explanation to a validated root cause before selecting countermeasures in Improve.

DMAIC phases with Analyse highlighted and the path from a candidate input to a validated root cause

A practical sequence is:

  1. Define the outcome, population, baseline, and improvement that matters.
  2. Confirm the measurement and choose the appropriate test.
  3. Write H₀ and H₁; set α and the one- or two-tailed direction in advance.
  4. Collect representative data and assess assumptions.
  5. Calculate the statistic and p-value, then interpret the decision alongside effect size and business relevance.

This connects hypothesis testing to Y = f(x): test whether a candidate input, X, meaningfully influences the process outcome, Y. A charter that cannot be translated into testable H₀ and H₁ by the end of Define may still describe an important topic, but it is not yet a project with a measurable, testable claim. For more on scoping, see Define-phase project scoping best practices.

The remaining rungs in this series

This first rung establishes the language of the test. The next guides will build toward practical application and more advanced decisions:

  • Rung 2: P-values and what they do, and do not, tell you.
  • Rung 3: Type I and Type II errors, sample size, and power.
  • Rung 4: One-sample, two-sample, and paired t-tests.
  • Rung 5: Variance pre-checks and choosing a suitable comparison.
  • Rung 6: ANOVA for comparing three or more means.
  • Rung 7: Chi-square tests and non-parametric alternatives.
  • Rung 8: Regression as hypothesis testing.
  • Rung 9: Equivalence testing when “no meaningful difference” is the claim.
  • Rung 10: Combining test choice, assumptions, and project decisions.

Build the capability to make evidence-based decisions

Hypotheses are more than two lines in a statistics worksheet. They define what a project can reasonably claim, guide analysis in DMAIC, and help leaders distinguish a credible improvement from an encouraging but uncertain pilot result.

Build your skills with our CSSC-accredited Lean Six Sigma Green Belt training or Black Belt online training. Both are self-paced and include practical learning through real-world simulations, dummy data, charts, worked examples, and a full end-to-end DMAIC case study. Explore our free White Belt starter and practice exams as your next step.

Kaizen. Kai-Care. Kai-Done. Lean Six Sigma

Related Posts