A process does not need a complex measurement system to produce valuable evidence. Sometimes the most important question is direct:
Did the unit pass, or did it fail?
That binary result is known as attribute data. It may appear simple, but when collected consistently and analysed with the right quality tools, it can reveal defect patterns, process instability and improvement opportunities with remarkable clarity.
In the realm of Lean Six Sigma, attribute data is especially useful when teams need to monitor the proportion of defective units, track service errors or verify whether an outcome meets a defined customer requirement. It supports tools such as p-charts, defect tracking, yield calculations and Pareto analysis.
This guide explains what attribute data is, how it differs from variable data, what is a p chart, and how a practical quality decision can be powered by nothing more than carefully defined Pass/Fail outcomes.
What Is Attribute Data?
Attribute data is qualitative, categorical data that classifies an observation into a defined category. The most common categories are:
- Pass or Fail
- Conforming or Nonconforming
- Accept or Reject
- Complete or Incomplete
- On time or Late
- Correct or Incorrect
- Present or Absent
Unlike variable data, attribute data does not record the precise measurement. It records whether a condition has been met.
For example, a logistics team may classify each delivery as:
- On time: delivered by the promised date
- Late: delivered after the promised date
A healthcare team may classify each patient record as:
- Complete: all required fields are present
- Incomplete: one or more required fields are missing
A manufacturing team may classify each component as:
- Pass: meets all acceptance criteria
- Fail: does not meet at least one criterion
The fundamental purpose of attribute data is to convert an operational requirement into a consistent decision rule. Once that rule is applied reliably, teams can measure the percentage of failures and monitor how that percentage changes over time.
Why Pass/Fail Data Is Valuable in Lean Six Sigma
Attribute data is often faster, less expensive and easier to collect than continuous measurements. It is particularly valuable when:
-
The customer requirement is categorical.
A shipment either arrived on time or it did not. -
The inspection method is inherently visual or judgment-based.
For example, a label may be classified as correctly positioned or incorrectly positioned. -
A go/no-go gauge is used.
The result is an acceptance decision rather than a precise measurement. -
The organisation needs an immediate operational signal.
A supervisor may need to know whether a batch can be released, quarantined or escalated. -
The process has a large volume of transactions.
Automated systems can classify thousands of records as compliant or noncompliant.
However, attribute data is only useful when the categories are clearly defined. If two inspectors interpret “acceptable appearance” differently, the data may reflect measurement system variation rather than actual process variation.
A strong operational definition should state:
- What is being inspected
- What constitutes a pass
- What constitutes a failure
- Who performs the inspection
- When the inspection occurs
- How ambiguous cases are handled
- How the result is recorded
What Is a P Chart?
If you are asking what is a p chart, the concise answer is this:
A p-chart is an attribute control chart used to monitor the proportion of defective units in a process over time.
The “p” represents proportion. For each subgroup, the calculation is:
[
p_i = \frac{\text{Number of defective units in subgroup } i}{\text{Total units inspected in subgroup } i}
]
A p-chart is appropriate when each unit is classified as defective or nondefective and the subgroup sample size may vary.
For example:
- Monday: 8 defective invoices out of 160
- Tuesday: 11 defective invoices out of 220
- Wednesday: 7 defective invoices out of 140
The number inspected changes each day, but the proportion defective can still be compared.

A p-chart typically includes:
- Centre line: the average proportion defective
- Upper control limit: the expected upper boundary of common-cause variation
- Lower control limit: the expected lower boundary, never less than zero
- Subgroup points: the observed proportion defective for each time period
For a constant subgroup size (n), the approximate control limits are:
[
\bar{p} = \frac{\text{Total defectives}}{\text{Total units inspected}}
]
[
UCL = \bar{p} + 3\sqrt{\frac{\bar{p}(1-\bar{p})}{n}}
]
[
LCL = \bar{p} – 3\sqrt{\frac{\bar{p}(1-\bar{p})}{n}}
]
If the subgroup sizes vary, the control limits should be calculated for each subgroup because the level of statistical precision changes with sample size.
A p-chart does not identify the root cause by itself. Instead, it signals where and when the team should investigate.
Worked Example: How Attribute Data Powered a Quality Decision
Consider a service centre processing insurance claims. Each claim is reviewed before release and classified as either:
- Pass: no required information is missing and the claim is processed correctly
- Fail: at least one preventable processing error is found
The team inspects 200 claims per day for 10 consecutive days. The results are:
| Day | Claims inspected | Failed claims | Proportion failed |
|---|---|---|---|
| 1 | 200 | 10 | 5.0% |
| 2 | 200 | 11 | 5.5% |
| 3 | 200 | 9 | 4.5% |
| 4 | 200 | 12 | 6.0% |
| 5 | 200 | 8 | 4.0% |
| 6 | 200 | 10 | 5.0% |
| 7 | 200 | 11 | 5.5% |
| 8 | 200 | 9 | 4.5% |
| 9 | 200 | 10 | 5.0% |
| 10 | 200 | 22 | 11.0% |
Across the 10 days, the team inspected 2,000 claims and identified 112 failures.
Therefore:
[
\bar{p} = \frac{112}{2000} = 0.056
]
The average failure proportion is 5.6%, equivalent to a baseline yield of 94.4%.
For (n=200):
- Centre line = 5.6%
- Approximate UCL = 10.5%
- LCL = 0%, because the calculated lower limit is negative
Day 10 produced an 11.0% failure rate, which exceeds the upper control limit. This is evidence of a likely special cause rather than ordinary process fluctuation.
The team does not immediately blame the claims processors or launch broad retraining. Instead, it stratifies the failed claims by:
- Processing shift
- Claim type
- Software version
- Processor experience
- Error category
The analysis shows that 18 of the 22 failures occurred on the evening shift, shortly after a workflow update changed the order of mandatory data fields. The corrective action is to restore the clearer field sequence, add an on-screen validation check and conduct a focused verification review.
During the following 10 days, the team inspects another 2,000 claims and records only 50 failures:
[
\text{Post-improvement failure rate} = \frac{50}{2000} = 2.5%
]
The results show:
- Failure rate reduced from 5.6% to 2.5%
- Relative reduction of approximately 55.4%
- Yield improved from 94.4% to 97.5%
- 62 fewer failed claims across 2,000 transactions
The Pass/Fail classification did not explain the cause. It did something equally important: it identified the signal that justified investigation and confirmed that the corrective action produced a meaningful improvement.
Attribute Data Versus Variable Data
The choice between attribute and variable data should be deliberate.
Attribute data
Attribute data is appropriate when the decision is naturally categorical or when the process requirement is expressed as acceptance or rejection.
Examples include:
- Whether a product passes a visual inspection
- Whether a customer request was resolved correctly
- Whether a delivery met its promised date
- Whether a medical record contains all required information
- Whether a transaction complied with a policy
Its strengths include speed, accessibility and easy communication. Its limitation is that it can hide important detail. A Pass/Fail result does not show how close a measurement was to its specification limit.
Variable data
Variable data records a measurable quantity on a continuous scale, such as:
- Processing time
- Weight
- Temperature
- Diameter
- Waiting time
- Fill volume
- Distance
- Response time
Suppose a packaging process must fill containers to 500 ± 5 grams. Attribute data can show whether each container passes or fails. Variable data can show whether the average is drifting toward 505 grams, whether the spread is widening and whether the process is centred within specifications.
Variable data is generally preferable when:
- The exact size of the variation matters
- Early detection of drift is important
- The specification has upper and lower limits
- The team needs detailed root-cause analysis
- The measurement can be collected reliably
Variable control charts such as X-bar and R charts or Individuals and Moving Range charts are designed for measured data. A p-chart should not be used as a substitute for a variable chart.
The practical rule is straightforward:
Use attribute data to classify outcomes. Use variable data to understand the magnitude of performance.
In mature Lean Six Sigma projects, teams may use both. A variable measurement can explain why units are failing, while attribute data confirms whether customers ultimately receive conforming output.
How Attribute Data Supports Defect Tracking and Yield
Attribute data is foundational to defect tracking, but practitioners must distinguish between a defective unit and a defect.
- A defective unit is an item that fails the overall acceptance requirement.
- A defect is a specific nonconformance found on a unit.
One invoice may be defective because it contains three separate errors. A p-chart tracks the proportion of defective invoices. A defect-counting chart, such as a c-chart or u-chart, may be more appropriate when the number of errors per invoice is the primary concern.
Attribute data also supports several important quality metrics:
First Pass Yield
First Pass Yield (FPY) is the percentage of units that pass a process step without rework, repair or correction.
[
FPY = \frac{\text{Units passing first time}}{\text{Total units entering the step}} \times 100
]
If 975 out of 1,000 units pass immediately, FPY is 97.5%.
Rolled Throughput Yield
Rolled Throughput Yield (RTY) multiplies the first-pass yields of several process steps. If three steps have yields of 98%, 97% and 99%:
[
RTY = 0.98 \times 0.97 \times 0.99 = 0.941
]
The overall rolled throughput yield is approximately 94.1%. This demonstrates why a high yield at each individual step can still produce a lower end-to-end result.
Defect Tracking
Teams can use attribute data to track:
- Defective units by day
- Failures by product or service type
- Error categories
- Defects by shift, location or team
- Customer complaints
- Rework and rejection rates
A Pareto chart can then rank failure categories, helping the project team focus on the most influential contributors.
Attribute Data, Zero Defects and Sustainable Control
The philosophy of Zero Defects, associated with Philip Crosby, promotes the principle of doing things right the first time. It does not mean assuming that every process will instantly achieve perfect performance. It means designing standards, controls and behaviours around prevention rather than accepting failure as an unavoidable cost of doing business.
Attribute data makes that philosophy visible. Every Pass/Fail result answers a practical question:
Did the process deliver what the customer, regulation or downstream operation required?
When connected to p-charts, yield measures, defect tracking and a disciplined response plan, attribute data becomes part of a wider control system. The objective is not to punish individuals for failures. It is to understand process performance and remove the conditions that make failure more likely.
Practical Rules for Using Attribute Data Well
Before building a p-chart or reporting a yield percentage, follow these protocols:
-
Define the unit clearly.
Decide whether the unit is an invoice, delivery, product, patient record or customer interaction. -
Create an operational definition.
Ensure different people classify the same outcome consistently. -
Record the denominator.
A failure count without the total inspected cannot produce a meaningful proportion. -
Separate defectives from defects.
Choose a p-chart when the focus is defective units, not the total number of defects. -
Use rational subgroups.
Organise observations by a meaningful period, shift, machine, product family or service channel. -
Investigate signals, not every fluctuation.
Look for points outside control limits, sustained runs and other non-random patterns. -
Recalculate responsibly.
After a confirmed process change, establish revised control limits using representative data. -
Connect quality data to customer requirements.
A lower failure rate matters most when it improves reliability, safety, timeliness or customer value.

Build the Capability to Turn Data Into Decisions
Attribute data is not a second-rate alternative to measurement data. It is the correct form of evidence when the process requirement is categorical, the inspection is pass/fail or the organisation needs a fast and reliable quality signal.
The real power comes from using it with discipline: define the categories, collect the denominator, select the right control chart, interpret variation correctly and connect the result to customer and business outcomes.
If you want to apply p-charts, yield analysis, defect tracking and the broader DMAIC framework with confidence, structured learning is the next step. Explore Lean Six Sigma certification or review the flexible online Lean Six Sigma training programmes offered by Lean 6 Sigma Hub. Our Green Belt certification develops the practical capability to analyse process data, lead improvement projects and sustain measurable gains.
Start your Lean Six Sigma training and pursue a recognised certification so you can turn everyday quality signals into confident, data-driven improvement decisions.
Kaizen. Kai-Care. Kai-Done. ( Lean Six Sigma)








