Statistics 2024 Paper I 50 marks Prove

Paper I — Q7

(a) A very big population is divided into two strata. The allocation of units of stratified random sample of size n for the two…

(a)

A very big population is divided into two strata. The allocation of units of stratified random sample of size n for the two strata under Neyman allocation are n'_1 and n'_2, and under other type of allocation are n_1 and n_2. Define r = n'_1/n'_2 and μ = n_1/(rn_2). Then prove that the efficiency of stratified random sampling with respect to stratified random sampling under Neyman allocation is given by e = μ(r+1)²/((μr + 1)(μ + r)). 20 marks

(b)
(i)

A bank has 40000 clients in its computer files, divided into 4000 branches, each managing exactly 10 clients. To estimate the proportion of clients for whom the bank has granted loan, a simple random sample of 40 branches is selected. From the selected sample, for each branch i, a list of clients (A_i) having a loan is prepared; i = 1, 2, ..., 40. The data observed from the selected sample are Σ(i=1 to 40) A_i = 200 and Σ(i=1 to 40) A_i² = 1156. What type of sampling is this? 3 marks

(ii)

State the expression of the parameter to estimate and obtain its unbiased estimate. 6 marks

(iii)

Estimate the variance of the unbiased estimator obtained in part (ii). 6 marks

(c)
(i)

Verify whether the following BIBD are possible: (1) v = b = 22, r = k = 7, λ = 2; (2) v = 10, b = 18, r = 9, k = 5, λ = 4. Given that the design is resolvable.

(ii)

Given below is the incidence matrix (N) of a block design. Find the degrees of freedom associated with the adjusted treatment sum of squares and the degrees of freedom for the error sum of squares.

हिंदी में प्रश्न पढ़ें
(a)

एक बहुत बड़ी समष्टि को दो स्तरों में विभाजित किया गया है। नेमन नियतन के अनुसार, दो स्तरों के लिए, आमाप n के स्तरीत यादृच्छिक प्रतिदर्श की इकाइयों के नियतन n'_1 और n'_2 हैं और दूसरे प्रकार की नियतन विधि के अनुसार n_1 तथा n_2 हैं। r = n'_1/n'_2 तथा μ = n_1/(rn_2) को परिभाषित कीजिए। तब सिद्ध कीजिए कि नेमन नियतन के अंतर्गत स्तरीत यादृच्छिक प्रतिचयन के सापेक्ष स्तरीत यादृच्छिक प्रतिचयन की दक्षता e = μ(r+1)²/((μr + 1)(μ + r)) है। (20 अंक)

(b)
(i)

एक बैंक में इसके कम्प्यूटर की फाइलों में 40000 ग्राहक हैं, जिनको 4000 शाखाओं में बाँटा गया है, प्रत्येक शाखा ठीक 10 ग्राहकों का प्रबंध करती है। जिन ग्राहकों को बैंक का ऋण दिया गया है, उनके अनुपात का आकलन करने के लिए 40 शाखाओं का एक सरल यादृच्छिक प्रतिदर्श चुना गया है। चुने गये प्रतिदर्श में से, प्रत्येक शाखा i के लिए ग्राहकों, जिन्होंने ऋण लिया है, उनकी एक सूची (A_i) तैयार की गई है; i = 1, 2, ..., 40 है। चयनित प्रतिदर्श से प्रेक्षित आँकड़े Σ(i=1 से 40) A_i = 200 और Σ(i=1 से 40) A_i² = 1156 प्राप्त हुए हैं। यह किस प्रकार का प्रतिचयन है? (3 अंक)

(ii)

प्राचल जिसका आकलन करना है, उसके लिए व्यंजक (एक्सप्रेशन) लिखिए और उसका अनभिनत आकलक ज्ञात कीजिए। (6 अंक)

(iii)

भाग (ii) में प्राप्त अनभिनत आकलक के विचरण का आकलन कीजिए। (6 अंक)

(c)
(i)

सत्यापित कीजिए कि क्या नीचे दिये गये BIBD संभव हैं: (1) v = b = 22, r = k = 7, λ = 2; (2) v = 10, b = 18, r = 9, k = 5, λ = 4। दिया गया है कि अभिकल्पना विभोज्य है।

(ii)

नीचे एक खंडक अभिकल्पना का आपतन आव्यूह (N) दिया गया है। समायोजित उपचार वर्गों के योग से संबद्ध स्वातंत्र्य-कोटियाँ और त्रुटि वर्गों के योग के लिए स्वातंत्र्य-कोटियाँ ज्ञात कीजिए।

Q7 of the 2024 UPSC Mains Statistics Paper I, as printed
The question as printed in the 2024 Statistics paper

The figure this question refers to, in words

The question paper is a scan and the diagram did not survive as text. This is the figure as read from the original page — every component, value and label — so the question can be worked from the text below.

(b) Summation from i=1 to 40 of A_i = 200 and summation from i=1 to 40 of A_i squared = 1156.

(c) Incidence matrix N of a block design, given as a 6x4 matrix with rows: row 1 = [1, 0, 0, 0]; row 2 = [1, 0, 1, 0]; row 3 = [1, 0, 1, 0]; row 4 = [0, 1, 0, 1]; row 5 = [0, 1, 0, 1]; row 6 = [0, 0, 0, 1]. The question asks to find the degrees of freedom associated with the adjusted treatment sum of squares and the degrees of freedom for the error sum of squares.

Model answer

Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.

(a) Let the two strata have sizes N₁, N₂, weights W₁ = N₁/N, W₂ = N₂/N and variances S₁², S₂². For a stratified random sample of size n, the variance of the unbiased estimator of the population mean is V = W₁²S₁²/n₁ + W₂²S₂²/n₂. Put A = W₁²S₁², B = W₂²S₂². Under Neyman allocation n′₁ ∝ W₁S₁ and n′₂ ∝ W₂S₂, so r = n′₁/n′₂ = W₁S₁/(W₂S₂), hence r² = A/B, i.e. B = A/r². Neyman allocation gives n′₁ = nr/(r+1), n′₂ = n/(r+1). Therefore V_N = A/n′₁ + B/n′₂ = (r+1)/n (A/r + B) = (r+1)/n (A/r + A/r²) = A(r+1)²/(n r²). For the other allocation, μ = n₁/(r n₂), so n₁ = μ r n/(μ r + 1), n₂ = n/(μ r + 1). Hence V = A/n₁ + B/n₂ = (μ r + 1)/n (A/(μ r) + B) = (μ r + 1)/n (A/(μ r) + A/r²) = A(μ r + 1)(μ + r)/(n μ r²). Efficiency of Neyman allocation relative to the other allocation is e = V_N/V = [A(r+1)²/(n r²)] / [A(μ r + 1)(μ + r)/(n μ r²)] = μ(r+1)²/((μ r + 1)(μ + r)). Hence proved.

(b)(i) This is single-stage cluster sampling with equal clusters. The branches are clusters, each of size 10. A simple random sample of 40 branches is selected, and all clients in the selected branches are examined.

(b)(ii) Let Aᵢ be the number of loan clients in branch i, i = 1, ..., 4000. The parameter is the population proportion P = (Σᵢ=1^4000 Aᵢ)/(4000 × 10) = bar A/10, where bar A is the population mean number of loan clients per branch. The unbiased estimator is hat P = (Σᵢ=1^40 Aᵢ)/(40 × 10) = 200/400 = 1/2 = 0.5.

(b)(iii) Here N = 4000, n = 40, M = 10. Sample mean per branch is bar A = 200/40 = 5. Sample variance of Aᵢ is s_A² = [Σ Aᵢ² - n(bar A)²]/(n - 1) = (1156 - 40 × 25)/39 = 156/39 = 4. Finite population correction f = n/N = 40/4000 = 0.01. Hence Var(hat P) = (1 - f) s_A²/(n M²) = (0.99 × 4)/(40 × 100) = 3.96/4000 = 0.00099. Thus the estimated variance is 0.00099 (proportion squared).

(c)(i) For a BIBD, necessary conditions are vr = bk and λ(v - 1) = r(k - 1). For a resolvable BIBD, v/k must be an integer. (1) v = b = 22, r = k = 7, λ = 2. The two equations hold: 22×7 = 22×7 and 2×21 = 7×6 = 42. But v/k = 22/7 is not an integer. Hence it is not possible as a resolvable BIBD. (Also, by Bruck-Ryser-Chowla, since v is even, k - λ = 5 is not a perfect square, so no symmetric BIBD exists.) (2) v = 10, b = 18, r = 9, k = 5, λ = 4. Here vr = 10×9 = 90 = 18×5, and λ(v - 1) = 4×9 = 36 = 9×4. Also v/k = 10/5 = 2, and b = r(v/k) = 9×2 = 18. However, a resolvable design with v = 2k and λ = k - 1 exists only if a Hadamard matrix of order v = 10 exists. Hadamard matrices exist only for orders 1, 2 or multiples of 4; 10 is not such an order. Hence this is not possible as a resolvable BIBD.

(c)(ii) The incidence matrix N is 6×4, so v = 6 treatments and b = 4 blocks. Total observations n = sum of entries = 1 + 2 + 2 + 2 + 2 + 1 = 10. The incidence graph has two connected components: treatments 1,2,3 with blocks 1,3, and treatments 4,5,6 with blocks 2,4. Thus c = 2. Adjusted treatment degrees of freedom = v - c = 6 - 2 = 4. Error degrees of freedom = n - b - v + c = 10 - 4 - 6 + 2 = 2.

What "Prove" is asking you to do

Establish that the statement holds for every case it claims, not for one representative case. The argument must be closed: each line follows from a definition, a hypothesis, or a named theorem you are entitled to use.

Structure that answers it

Given and to prove, restated → theorem or construction to be used, named → the argument line by line → conclusion stated as proved

Where marks are lost

Testing one example, which illustrates but proves nothing. On an if and only if claim, proving one direction and stopping forfeits that half outright, and degenerate cases — zero, the empty set, the equality case — have to be disposed of rather than assumed away.

All UPSC directive words, compared →

How this answer will be evaluated

Approach

Framework: Stratified Sampling Efficiency & Cluster Sampling Variance. (a) derive: given > assumptions > stepwise derivation > result > check | (b) explain: definition/context > points in order > small example > short close | (c) derive: given > assumptions > stepwise derivation > result > check Full marks: Rigorous derivation with all steps shown; correct identification of sampling type; accurate variance calculation; proper BIBD verification and df calculation.

Key points expected

  • Efficiency formula derivation with r and μ substitution
  • One-stage cluster sampling identification
  • Proportion estimation p_hat = 0.5
  • Cluster variance calculation using ΣA_i and ΣA_i²
  • BIBD condition verification for both cases
  • Resolvable design check
  • Rank calculation of N'N for degrees of freedom

Evaluation rubric

Each sub-part is marked on its own, against the marks and word limit printed on the paper.

  1. (a) Derive the efficiency formula for stratified sampling relative to Neyman allocation. 20 marks

    derive— given → assumptions → stepwise derivation → result → check

    Must cover

    • State variance formula for general allocation
    • State variance formula for Neyman allocation
    • Substitute r and μ definitions correctly
    • Simplify to required e = μ(r+1)²/((μr+1)(μ+r))

    Loses marks

    • Skipping substitution of r and μ
    • Algebraic errors in simplification

    Earns more

    • Explicitly define stratum weights W_h
    • Show intermediate algebraic steps clearly

    Extra mark

    • Mention condition for maximum efficiency (μ=1)
  2. (b) Identify sampling type, estimate parameter, and calculate variance. 15 marks

    explain— definition/context → points in order → small example → short close

    Must cover

    • Identify as One-Stage Cluster Sampling
    • Estimate proportion p = ΣA_i / (n*M)
    • Calculate p_hat = 200 / (40*10) = 0.5
    • Use cluster variance formula for estimation

    Loses marks

    • Confusing cluster sampling with stratified
    • Using simple random sampling variance formula

    Earns more

    • Correctly identify M=10 as cluster size
    • Show calculation of sample variance s_a²

    Extra mark

    • Mention finite population correction if applicable
  3. (c) Verify BIBD existence and calculate degrees of freedom from incidence matrix. 15 marks

    derive— given → assumptions → stepwise derivation → result → check

    Must cover

    • Check BIBD conditions: vr=bk and λ(v-1)=r(k-1)
    • Verify resolvable condition (r divisible by number of blocks)
    • Calculate rank of N'N for treatment df
    • Calculate error df = v - rank(N'N)

    Loses marks

    • Failing to check both BIBD conditions
    • Incorrect rank calculation for N'N

    Earns more

    • Explicitly show matrix multiplication for N'N
    • State that design is resolvable implies balanced incomplete block design

    Extra mark

    • Mention specific BIBD parameters for case (1)

Practice this exact question

Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.

Evaluate my answer →

More from Statistics 2024 Paper I