Paper I — Q7
(a) A very big population is divided into two strata. The allocation of units of stratified random sample of size n for the two…
A very big population is divided into two strata. The allocation of units of stratified random sample of size n for the two strata under Neyman allocation are n'_1 and n'_2, and under other type of allocation are n_1 and n_2. Define r = n'_1/n'_2 and μ = n_1/(rn_2). Then prove that the efficiency of stratified random sampling with respect to stratified random sampling under Neyman allocation is given by e = μ(r+1)²/((μr + 1)(μ + r)). 20 marks
A bank has 40000 clients in its computer files, divided into 4000 branches, each managing exactly 10 clients. To estimate the proportion of clients for whom the bank has granted loan, a simple random sample of 40 branches is selected. From the selected sample, for each branch i, a list of clients (A_i) having a loan is prepared; i = 1, 2, ..., 40. The data observed from the selected sample are Σ(i=1 to 40) A_i = 200 and Σ(i=1 to 40) A_i² = 1156. What type of sampling is this? 3 marks
State the expression of the parameter to estimate and obtain its unbiased estimate. 6 marks
Estimate the variance of the unbiased estimator obtained in part (ii). 6 marks
Verify whether the following BIBD are possible: (1) v = b = 22, r = k = 7, λ = 2; (2) v = 10, b = 18, r = 9, k = 5, λ = 4. Given that the design is resolvable.
Given below is the incidence matrix (N) of a block design. Find the degrees of freedom associated with the adjusted treatment sum of squares and the degrees of freedom for the error sum of squares.
हिंदी में प्रश्न पढ़ें
एक बहुत बड़ी समष्टि को दो स्तरों में विभाजित किया गया है। नेमन नियतन के अनुसार, दो स्तरों के लिए, आमाप n के स्तरीत यादृच्छिक प्रतिदर्श की इकाइयों के नियतन n'_1 और n'_2 हैं और दूसरे प्रकार की नियतन विधि के अनुसार n_1 तथा n_2 हैं। r = n'_1/n'_2 तथा μ = n_1/(rn_2) को परिभाषित कीजिए। तब सिद्ध कीजिए कि नेमन नियतन के अंतर्गत स्तरीत यादृच्छिक प्रतिचयन के सापेक्ष स्तरीत यादृच्छिक प्रतिचयन की दक्षता e = μ(r+1)²/((μr + 1)(μ + r)) है। (20 अंक)
एक बैंक में इसके कम्प्यूटर की फाइलों में 40000 ग्राहक हैं, जिनको 4000 शाखाओं में बाँटा गया है, प्रत्येक शाखा ठीक 10 ग्राहकों का प्रबंध करती है। जिन ग्राहकों को बैंक का ऋण दिया गया है, उनके अनुपात का आकलन करने के लिए 40 शाखाओं का एक सरल यादृच्छिक प्रतिदर्श चुना गया है। चुने गये प्रतिदर्श में से, प्रत्येक शाखा i के लिए ग्राहकों, जिन्होंने ऋण लिया है, उनकी एक सूची (A_i) तैयार की गई है; i = 1, 2, ..., 40 है। चयनित प्रतिदर्श से प्रेक्षित आँकड़े Σ(i=1 से 40) A_i = 200 और Σ(i=1 से 40) A_i² = 1156 प्राप्त हुए हैं। यह किस प्रकार का प्रतिचयन है? (3 अंक)
प्राचल जिसका आकलन करना है, उसके लिए व्यंजक (एक्सप्रेशन) लिखिए और उसका अनभिनत आकलक ज्ञात कीजिए। (6 अंक)
भाग (ii) में प्राप्त अनभिनत आकलक के विचरण का आकलन कीजिए। (6 अंक)
सत्यापित कीजिए कि क्या नीचे दिये गये BIBD संभव हैं: (1) v = b = 22, r = k = 7, λ = 2; (2) v = 10, b = 18, r = 9, k = 5, λ = 4। दिया गया है कि अभिकल्पना विभोज्य है।
नीचे एक खंडक अभिकल्पना का आपतन आव्यूह (N) दिया गया है। समायोजित उपचार वर्गों के योग से संबद्ध स्वातंत्र्य-कोटियाँ और त्रुटि वर्गों के योग के लिए स्वातंत्र्य-कोटियाँ ज्ञात कीजिए।
The figure this question refers to, in words
The question paper is a scan and the diagram did not survive as text. This is the figure as read from the original page — every component, value and label — so the question can be worked from the text below.
(b) Summation from i=1 to 40 of A_i = 200 and summation from i=1 to 40 of A_i squared = 1156.
(c) Incidence matrix N of a block design, given as a 6x4 matrix with rows: row 1 = [1, 0, 0, 0]; row 2 = [1, 0, 1, 0]; row 3 = [1, 0, 1, 0]; row 4 = [0, 1, 0, 1]; row 5 = [0, 1, 0, 1]; row 6 = [0, 0, 0, 1]. The question asks to find the degrees of freedom associated with the adjusted treatment sum of squares and the degrees of freedom for the error sum of squares.
Model answer
Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.
(a) Let the two strata have sizes N₁, N₂, weights W₁ = N₁/N, W₂ = N₂/N and variances S₁², S₂². For a stratified random sample of size n, the variance of the unbiased estimator of the population mean is V = W₁²S₁²/n₁ + W₂²S₂²/n₂. Put A = W₁²S₁², B = W₂²S₂². Under Neyman allocation n′₁ ∝ W₁S₁ and n′₂ ∝ W₂S₂, so r = n′₁/n′₂ = W₁S₁/(W₂S₂), hence r² = A/B, i.e. B = A/r². Neyman allocation gives n′₁ = nr/(r+1), n′₂ = n/(r+1). Therefore V_N = A/n′₁ + B/n′₂ = (r+1)/n (A/r + B) = (r+1)/n (A/r + A/r²) = A(r+1)²/(n r²). For the other allocation, μ = n₁/(r n₂), so n₁ = μ r n/(μ r + 1), n₂ = n/(μ r + 1). Hence V = A/n₁ + B/n₂ = (μ r + 1)/n (A/(μ r) + B) = (μ r + 1)/n (A/(μ r) + A/r²) = A(μ r + 1)(μ + r)/(n μ r²). Efficiency of Neyman allocation relative to the other allocation is e = V_N/V = [A(r+1)²/(n r²)] / [A(μ r + 1)(μ + r)/(n μ r²)] = μ(r+1)²/((μ r + 1)(μ + r)). Hence proved.
(b)(i) This is single-stage cluster sampling with equal clusters. The branches are clusters, each of size 10. A simple random sample of 40 branches is selected, and all clients in the selected branches are examined.
(b)(ii) Let Aᵢ be the number of loan clients in branch i, i = 1, ..., 4000. The parameter is the population proportion P = (Σᵢ=1^4000 Aᵢ)/(4000 × 10) = bar A/10, where bar A is the population mean number of loan clients per branch. The unbiased estimator is hat P = (Σᵢ=1^40 Aᵢ)/(40 × 10) = 200/400 = 1/2 = 0.5.
(b)(iii) Here N = 4000, n = 40, M = 10. Sample mean per branch is bar A = 200/40 = 5. Sample variance of Aᵢ is s_A² = [Σ Aᵢ² - n(bar A)²]/(n - 1) = (1156 - 40 × 25)/39 = 156/39 = 4. Finite population correction f = n/N = 40/4000 = 0.01. Hence Var(hat P) = (1 - f) s_A²/(n M²) = (0.99 × 4)/(40 × 100) = 3.96/4000 = 0.00099. Thus the estimated variance is 0.00099 (proportion squared).
(c)(i) For a BIBD, necessary conditions are vr = bk and λ(v - 1) = r(k - 1). For a resolvable BIBD, v/k must be an integer. (1) v = b = 22, r = k = 7, λ = 2. The two equations hold: 22×7 = 22×7 and 2×21 = 7×6 = 42. But v/k = 22/7 is not an integer. Hence it is not possible as a resolvable BIBD. (Also, by Bruck-Ryser-Chowla, since v is even, k - λ = 5 is not a perfect square, so no symmetric BIBD exists.) (2) v = 10, b = 18, r = 9, k = 5, λ = 4. Here vr = 10×9 = 90 = 18×5, and λ(v - 1) = 4×9 = 36 = 9×4. Also v/k = 10/5 = 2, and b = r(v/k) = 9×2 = 18. However, a resolvable design with v = 2k and λ = k - 1 exists only if a Hadamard matrix of order v = 10 exists. Hadamard matrices exist only for orders 1, 2 or multiples of 4; 10 is not such an order. Hence this is not possible as a resolvable BIBD.
(c)(ii) The incidence matrix N is 6×4, so v = 6 treatments and b = 4 blocks. Total observations n = sum of entries = 1 + 2 + 2 + 2 + 2 + 1 = 10. The incidence graph has two connected components: treatments 1,2,3 with blocks 1,3, and treatments 4,5,6 with blocks 2,4. Thus c = 2. Adjusted treatment degrees of freedom = v - c = 6 - 2 = 4. Error degrees of freedom = n - b - v + c = 10 - 4 - 6 + 2 = 2.
What "Prove" is asking you to do
Establish that the statement holds for every case it claims, not for one representative case. The argument must be closed: each line follows from a definition, a hypothesis, or a named theorem you are entitled to use.
Structure that answers it
Given and to prove, restated → theorem or construction to be used, named → the argument line by line → conclusion stated as proved
Where marks are lost
Testing one example, which illustrates but proves nothing. On an if and only if claim, proving one direction and stopping forfeits that half outright, and degenerate cases — zero, the empty set, the equality case — have to be disposed of rather than assumed away.
How this answer will be evaluated
Approach
Framework: Stratified Sampling Efficiency & Cluster Sampling Variance. (a) derive: given > assumptions > stepwise derivation > result > check | (b) explain: definition/context > points in order > small example > short close | (c) derive: given > assumptions > stepwise derivation > result > check Full marks: Rigorous derivation with all steps shown; correct identification of sampling type; accurate variance calculation; proper BIBD verification and df calculation.
Key points expected
- Efficiency formula derivation with r and μ substitution
- One-stage cluster sampling identification
- Proportion estimation p_hat = 0.5
- Cluster variance calculation using ΣA_i and ΣA_i²
- BIBD condition verification for both cases
- Resolvable design check
- Rank calculation of N'N for degrees of freedom
Evaluation rubric
Each sub-part is marked on its own, against the marks and word limit printed on the paper.
- (a) Derive the efficiency formula for stratified sampling relative to Neyman allocation. 20 marks
derive— given → assumptions → stepwise derivation → result → check
Must cover
- State variance formula for general allocation
- State variance formula for Neyman allocation
- Substitute r and μ definitions correctly
- Simplify to required e = μ(r+1)²/((μr+1)(μ+r))
Loses marks
- Skipping substitution of r and μ
- Algebraic errors in simplification
Earns more
- Explicitly define stratum weights W_h
- Show intermediate algebraic steps clearly
Extra mark
- Mention condition for maximum efficiency (μ=1)
- (b) Identify sampling type, estimate parameter, and calculate variance. 15 marks
explain— definition/context → points in order → small example → short close
Must cover
- Identify as One-Stage Cluster Sampling
- Estimate proportion p = ΣA_i / (n*M)
- Calculate p_hat = 200 / (40*10) = 0.5
- Use cluster variance formula for estimation
Loses marks
- Confusing cluster sampling with stratified
- Using simple random sampling variance formula
Earns more
- Correctly identify M=10 as cluster size
- Show calculation of sample variance s_a²
Extra mark
- Mention finite population correction if applicable
- (c) Verify BIBD existence and calculate degrees of freedom from incidence matrix. 15 marks
derive— given → assumptions → stepwise derivation → result → check
Must cover
- Check BIBD conditions: vr=bk and λ(v-1)=r(k-1)
- Verify resolvable condition (r divisible by number of blocks)
- Calculate rank of N'N for treatment df
- Calculate error df = v - rank(N'N)
Loses marks
- Failing to check both BIBD conditions
- Incorrect rank calculation for N'N
Earns more
- Explicitly show matrix multiplication for N'N
- State that design is resolvable implies balanced incomplete block design
Extra mark
- Mention specific BIBD parameters for case (1)
Practice this exact question
Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.
Evaluate my answer →More from Statistics 2024 Paper I
- Q4 (a) Find the most powerful test of size α(= 0·05) for testing H₀: μ = 0 vs. H₁: μ = 1, gi…
- Q5 (a) How will you justify the usage of the principle of least squares in estimating the pa…
- Q6 (a) (X, Y) has bivariate normal distribution BN(μ₁, μ₂, σ₁², σ₂², ρ). (i) Show that X and…
- Q7 (a) A very big population is divided into two strata. The allocation of units of stratifi…
- Q8 (a) A 2²-factorial design was used to develop the yield of a crop. Two factors A and B we…