Paper I — Q7
(a) Discuss the difference between sampling for variables and sampling for attributes with examples. For a qualitative…
Discuss the difference between sampling for variables and sampling for attributes with examples. For a qualitative characteristic, find an unbiased estimator of population proportion along with its variance when sample is drawn by simple random sampling without replacement. Also obtain an unbiased estimator of this variance. 20 marks
The table given below gives the population and sample sizes, stratum means and variance of a stratified random sample of size 50. Symbols used have their usual meanings.
| Stratum Number | Nᵢ | nᵢ | ȳᵢ | sᵢ² |
|---|---|---|---|---|
| 1 | 30 | 5 | 35 | 36 |
| 2 | 50 | 10 | 40 | 49 |
| 3 | 60 | 15 | 40 | 81 |
| 4 | 60 | 20 | 55 | 144 |
Verify that the existing allocation is optimum for given 4 strata. Also calculate the estimate of population variance under this allocation. 15 marks
Differentiate between Simple Random Sampling and Probability Proportional to Size Sampling. How will you draw a PPS sample of size n from a population of size N (n < N) by (i) Cumulative Total Method and (ii) Lahri's Method ? Explain. 15 marks
हिंदी में प्रश्न पढ़ें
चरों के प्रतिचयन एवं गुणात्मक चरों के प्रतिचयन में अंतर का उदाहरणों सहित वर्णन कीजिए। एक गुणात्मक अभिलक्षण के लिए समष्टि अनुपात का अनभिनत आकलक तथा इस आकलक का प्रसरण ज्ञात कीजिए जबकि प्रतिचयन प्रतिस्थापन रहित सरल यादृच्छिक विधि द्वारा किया गया है। इस प्रसरण का अनभिनत आकलक भी निकालिए। 20
नीचे दी गई सारणी में 50 आकार के स्तरीकृत यादृच्छिक प्रतिदर्श के स्तरों का माध्य एवं प्रसरण तथा स्तरों की समष्टि का आकार तथा स्तरों से चयनित प्रतिदर्शी आकारों को दिया गया है। चिह्नों को उनके सामान्य अर्थों में प्रयुक्त किया गया है।
| स्तर संख्या | Nᵢ | nᵢ | ȳᵢ | sᵢ² |
|---|---|---|---|---|
| 1 | 30 | 5 | 35 | 36 |
| 2 | 50 | 10 | 40 | 49 |
| 3 | 60 | 15 | 40 | 81 |
| 4 | 60 | 20 | 55 | 144 |
प्रमाणित कीजिए कि दिए गए 4 स्तरों के लिए मौजूदा आवंटन इष्टतम है। समष्टि प्रसरण का इस आवंटन के सापेक्ष आकलक भी ज्ञात कीजिए। 15
सरल यादृच्छिक प्रतिचयन तथा आकार अनुपातिक प्रायिकता प्रतिचयन में विभेद कीजिए । एक n आकार के आकार अनुपातिक प्रायिकता प्रतिदर्श को आप (i) संचयी योग विधि तथा (ii) लाहिरी विधि द्वारा N (n < N) आकार की समष्टि से कैसे चुनेंगे ? स्पष्ट कीजिए । 15
The figure this question refers to, in words
The question paper is a scan and the diagram did not survive as text. This is the figure as read from the original page — every component, value and label — so the question can be worked from the text below.
(b) Table with 5 rows and 5 columns. Header row: Stratum Number, N_i, n_i, y_bar_i, s_i^2. Row 1: 1, 30, 5, 35, 36. Row 2: 2, 50, 10, 40, 49. Row 3: 3, 60, 15, 40, 81. Row 4: 4, 60, 20, 55, 144.
Model answer
Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.
Sampling for variables and attributes. Sampling for variables treats the study characteristic as quantitative, so each unit has a measured value and the target is a mean or total. The distinction is practical: variable sampling needs measurement scales, while attribute sampling needs clear classification rules. Indian examples are per-hectare wheat yield, household monthly expenditure measured by NSSO, or factory output. Sampling for attributes treats the characteristic as qualitative or dichotomous, so each unit is classified as possessing or not possessing the attribute and the target is a proportion. Examples are whether a household is literate in the Census of India, whether a farm has irrigation, or whether a worker is employed. Variable sampling usually estimates ȳ or total Nȳ, while attribute sampling estimates P or NP. Let the population contain N units, M of which possess the attribute, so P=M/N. In a simple random sample without replacement of size n, let n' be the number possessing the attribute. Then p̂=n'/n, and E(p̂)=P, so p̂ is unbiased. For the binary variable yᵢ=1 if the unit has the attribute and 0 otherwise, the population variance is S²=N/(N-1)P(1-P). Therefore V(p̂)=(1-f)S²/n=(N-n)/(N-1)·P(1-P)/n. Since the sample variance s²=n/(n-1)p̂(1-p̂) is unbiased for S², an unbiased estimator of V(p̂) is v(p̂)=(N-n)/(N n)s²=(N-n)/(N(n-1))p̂(1-p̂).
Optimum allocation and variance. For equal costs, Neyman allocation requires nᵢ/n=N_iSᵢ/ΣN_jSⱼ; if costs differ, the condition becomes nᵢ∝N_iSᵢ/√cᵢ. Here Sᵢ=√(sᵢ²), so Sᵢ are 6, 7, 9 and 12. The products N_iSᵢ are 180, 350, 540 and 720, with total 1790. The optimum sample sizes are 50×180/1790≈5.03, 50×350/1790≈9.78, 50×540/1790≈15.08 and 50×720/1790≈20.11, which round to 5, 10, 15 and 20. The ratios nᵢ/(N_iSᵢ) are also nearly equal, confirming the given allocation as optimum. The stratum weights are Wᵢ=Nᵢ/200, namely 0.15, 0.25, 0.30 and 0.30. The stratified mean is ȳ_st=ΣWᵢȳᵢ=0.15×35+0.25×40+0.30×40+0.30×55=43.75. The estimated variance of this mean under the allocation is V̂(ȳ_st)=ΣWᵢ²(Nᵢ-nᵢ)/(Nᵢ nᵢ)sᵢ². The four terms are 0.15²×25/150×36=0.135, 0.25²×40/500×49=0.245, 0.30²×45/900×81=0.3645 and 0.30²×40/1200×144=0.432. Hence V̂(ȳ_st)=1.1765. This estimate uses the stratum sample variances and finite population corrections, and gives the standard error basis for ȳ_st.
SRS and PPS. Simple random sampling gives every unit the same probability of selection; it is easy to implement and unbiased, but it can be inefficient when units differ greatly in size or contribution. SRS is appropriate when the frame is homogeneous; PPS is appropriate when the frame is of establishments and size is known. PPS sampling gives each unit a selection probability proportional to an auxiliary size measure Xᵢ, such as area under crop, registered capital, or number of workers. It is therefore suitable for skewed populations such as large industrial units or large agricultural holdings, where a few large units dominate the total. For the usual PPS with replacement, let X=ΣXᵢ. In the cumulative total method, form cumulative totals Cᵢ=Σⱼ₌₁ⁱ Xⱼ. Draw n independent random numbers R₁,...,Rₙ from 1 to X; for each Rₖ select the unit i satisfying Cᵢ₋₁<Rₖ≤Cᵢ. Repetitions are allowed, so the same unit may be selected more than once. In Lahiri’s method, let M=max Xᵢ. Draw a random unit index i from 1 to N and a random number j from 1 to M; if j≤Xᵢ accept unit i, otherwise reject the pair and draw again. Conditional on acceptance, the probability of selecting unit i is Xᵢ/X, so each accepted draw is PPS. Repeat independently n times to obtain n draws. If a PPS without replacement sample is required, one should use systematic PPSWOR or a sequential PPSWOR method rather than merely discarding repetitions. Thus, the choice among SRS, stratified optimum allocation and PPS depends on whether the variable is quantitative or qualitative, how strata differ, and whether size measures explain variability.
What "Discuss" is asking you to do
Lay the issue out from more than one side — how it arose, what is claimed for it, what is held against it, and where it now stands. UPSC attaches discuss to broad topics with several live dimensions, so coverage of the dimensions earns more than the strength of your opinion.
Structure that answers it
Set the issue up → the case as it is made → the case against → the dimension both sides leave out → where the balance now lies
Where marks are lost
Listing facts with no thread between them, or arguing one side throughout and calling it a discussion.
How this answer will be evaluated
Approach
Framework: UPSC Statistics Paper 1. (a) discuss: intro > 3-4 dimensions > example > balanced close | (b) calculate: given > formula > substitution > result with units > interpretation | (c) explain: definition/context > points in order > small example > short close Full marks: All derivations complete, calculations accurate, methods clearly explained with correct notation
Key points expected
- Define sampling for variables vs attributes with examples
- Derive unbiased estimator of population proportion p
- Derive variance of estimator under SRSWOR
- Provide unbiased estimator of the variance
- Verify existing allocation is optimum (Neyman allocation)
- Calculate estimate of population variance
- Show calculation steps clearly
- Use correct stratified sampling formulas
Evaluation rubric
Each sub-part is marked on its own, against the marks and word limit printed on the paper.
- (a) Distinguish variable vs attribute sampling and derive unbiased estimator of proportion with variance. 20 marks
discuss— intro → 3-4 dimensions → example → balanced close
Must cover
- Define sampling for variables vs attributes with examples
- Derive unbiased estimator of population proportion p
- Derive variance of estimator under SRSWOR
- Provide unbiased estimator of the variance
Loses marks
- Confusing variable and attribute sampling
- Missing variance derivation steps
- No unbiased estimator of variance provided
Earns more
- Correct notation for estimator vs parameter
- Explicit statement of SRSWOR assumption
Extra mark
- Mention of finite population correction factor
- (b) Verify optimum allocation and calculate population variance estimate. 15 marks
calculate— given → formula → substitution → result with units → interpretation
Must cover
- Verify existing allocation is optimum (Neyman allocation)
- Calculate estimate of population variance
- Show calculation steps clearly
- Use correct stratified sampling formulas
Loses marks
- Incorrect allocation verification
- Arithmetic errors in variance calculation
- Missing formula before substitution
Earns more
- Clean table of intermediate calculations
- Correct interpretation of variance estimate
Extra mark
- Comparison with proportional allocation
- (c) Differentiate SRS and PPS sampling; explain PPS sample drawing methods. 15 marks
explain— definition/context → points in order → small example → short close
Must cover
- Differentiate SRS and PPS sampling
- Explain Cumulative Total Method for PPS
- Explain Lahri's Method for PPS
- Show how to draw PPS sample of size n
Loses marks
- Confusing SRS and PPS concepts
- Incomplete explanation of either method
- No clear procedure for sample drawing
Earns more
- Clear step-by-step procedure for each method
- Small numerical example for illustration
Extra mark
- Mention of when PPS is preferred over SRS
Practice this exact question
Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.
Evaluate my answer →More from Statistics 2023 Paper I
- Q4 (a) What is the role of properties of completeness and sufficiency in Statistical Inferen…
- Q5 (a) (i) If **X** = (X₁ X₂ X₃)' is distributed as N₃ (μ, Σ), find the distribution of [(X₁…
- Q6 (a) Let **X** = (X₁ X₂ X₃)' be distributed as N₃ (μ, Σ), where μ = (2 −1 3)' and Σ = 4 &…
- Q7 (a) Discuss the difference between sampling for variables and sampling for attributes wit…
- Q8 (a) Differentiate between randomised block design and balanced incomplete block design. I…