Practice a College Board-style free response question on describing a distribution, identifying outliers, and evaluating a sampling method. Write your response, then reveal the model answer.
A researcher wants to study how many hours per week students at a large high school spend on homework. She records the values for a group of students and computes the five-number summary (in hours) shown below. The distribution has a mean of about 10.5 hours.
| Statistic | Value (hours) |
|---|---|
| Minimum | 2 |
| Q1 | 6 |
| Median | 9 |
| Q3 | 13 |
| Maximum | 28 |
The distribution appears skewed to the right: the maximum (28) is much farther above Q3 than the minimum (2) is below Q1, and the mean (10.5) is greater than the median (9), which is typical of right skew. A reasonable measure of center is the median, 9 hours. For spread, the IQR = Q3 − Q1 = 13 − 6 = 7 hours (and the range is 28 − 2 = 26 hours).
IQR = 13 − 6 = 7, so 1.5 × IQR = 10.5. The upper fence is Q3 + 1.5·IQR = 13 + 10.5 = 23.5 hours. Since 28 > 23.5, the maximum value of 28 hours is an outlier. (The lower fence is 6 − 10.5 = −4.5, so the minimum of 2 is not a low outlier.)
This is a convenience sample — she chose students who were easy to reach rather than selecting randomly. A likely source of bias is undercoverage: students who spend time in the library after school are probably more studious than the overall student body, and students who never use the library have no chance of being selected. As a result, her sample would tend to include heavier homework-doers, so her estimate of the average homework time would most likely be biased too high — an overestimate — and it should not be generalized to all students at the school.