Statistics 2021 Paper I 50 marks Prove

Paper I — Q4

(a) Let X₁, X₂, ..., Xₙ be a random sample from Poisson distribution with mean λ > 0. Define a statistic W = (1 − 1/n)^T, T =…

(a)

Let X₁, X₂, ..., Xₙ be a random sample from Poisson distribution with mean λ > 0. Define a statistic W = (1 − 1/n)^T, T = Σᵢ₌₁ⁿ Xᵢ (i) Show that T is complete sufficient statistic. (ii) Show that T is unbiased for e^(−λ). (iii) Show that even though T is UMVUE, it does not attain the CRLB for g(λ) = e^(−λ). 20 marks

(b)

Let f(x, y) = (e^(-yx²)/2 y³/2 e^-y)/(√(2π)), -∞ < x < ∞, y > 0. (i) Obtain the marginal distribution of Y and conditional distribution of X given Y. (ii) Find E(Y), V(Y), E(X|Y), V(X|Y). (iii) Use (ii) to find E(X), V(X). (5+5+5 marks)

(c)

A company's trainees are randomly assigned to groups which are through a certain industrial inspection procedure by three different methods. At the end of the instructing period they are tested for inspection performance quality. The following are their scores : Method A : 80 83 79 85 90 68 Method B : 82 84 60 72 86 67 91 Method C : 93 65 77 78 88 Using the appropriate non-parametric test, determine at 0·05 level of significance whether the three methods are equally effective. 15 marks

हिंदी में प्रश्न पढ़ें
(a)

माना X₁, X₂, ..., Xₙ प्वासों बंटन, जिसका माध्य λ > 0, से लिया गया एक यादृच्छिक प्रतिदर्श है । एक प्रतिदर्शज परिभाषित है W = (1 − 1/n)^T, T = Σᵢ₌₁ⁿ Xᵢ (i) दर्शाइए कि T पूर्ण पर्याप्त प्रतिदर्शज है । (ii) दर्शाइए कि T, e^(−λ) के लिए अनभिनत है । (iii) भले ही T, UMVUE (यू.एम.वी.यू.ई.) है, दर्शाइए कि यह g(λ) = e^(−λ) के लिए CRLB (सी.आर.एल.बी.) प्राप्त नहीं करता है । (20 अंक)

(b)

माना f(x, y) = [e^(−yx²/2) y^(3/2) e^(−y)] / √(2π) , −∞ < x < ∞, y > 0. (i) Y का उपांत बंटन और Y के दिए होने पर X का सप्रतिबंध बंटन प्राप्त कीजिए । (ii) E(Y), V(Y), E(X|Y), V(X|Y) ज्ञात कीजिए । (iii) (ii) का उपयोग करते हुए E(X), V(X) ज्ञात कीजिए । (5+5+5 अंक)

(c)

तीन विभिन्न तरीकों से एक निश्चित औद्योगिक निरीक्षण प्रक्रिया द्वारा एक कंपनी के प्रशिक्षार्थियों को यादृच्छया समूहों में नियत किया गया । प्रशिक्षण अवधि की समाप्ति पर निरीक्षण प्रदर्शन गुणवत्ता के लिए उनका परीक्षण किया गया । उनके स्कोर निम्न हैं : रीति A : 80 83 79 85 90 68 रीति B : 82 84 60 72 86 67 91 रीति C : 93 65 77 78 88 उपयुक्त अप्राचलिक परीक्षण का उपयोग करते हुए, 0·05 सार्थकता स्तर पर निर्धारित कीजिए कि क्या तीनों रीतियाँ समान रूप से प्रभावी हैं । (15 अंक)

Q4 of the 2021 UPSC Mains Statistics Paper I, as printed
The question as printed in the 2021 Statistics paper

Model answer

Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.

(a)(i) The joint pmf is p(x₁,...,xₙ; λ) = ∏ᵢ e^(−λ) λ^(xᵢ)/xᵢ! = e^(−nλ) λ^(Σxᵢ)/∏xᵢ!.

By the Neyman factorization theorem, this factors as c(λ) λ^(T) h(x), so T = Σᵢ Xᵢ is sufficient for λ.

For completeness, T ~ Poisson(nλ), since the sum of independent Poisson variables is Poisson. If E_λ[h(T)] = 0 for all λ > 0, then Σₜ h(t) e^(−nλ)(nλ)^t/t! = 0. Put z = nλ > 0. Then Σₜ h(t) z^t/t! = 0 for all z > 0. A power series identically zero must have all coefficients zero, so h(t) = 0 for every t. Hence T is complete sufficient.

(ii) As stated, T itself is not unbiased for e^(−λ), because E(T) = nλ. The intended unbiased statistic is W = (1 − 1/n)^T. Using the Poisson pgf, E(W) = E[(1 − 1/n)^T] = Σₜ (1 − 1/n)^t e^(−nλ)(nλ)^t/t! = e^(−nλ) exp[nλ(1 − 1/n)] = e^(−nλ) e^(nλ−λ) = e^(−λ). Thus W is unbiased for e^(−λ).

(iii) Since T is complete sufficient and W is a function of T unbiased for e^(−λ), by the Lehmann–Scheffé theorem W is the UMVUE of e^(−λ).

For X ~ Poisson(λ), log f = x log λ − λ − log x!, so the score is U = x/λ − 1. Hence I₁(λ) = Var(X/λ − 1) = Var(X)/λ² = λ/λ² = 1/λ. For n observations, I_n(λ) = n/λ. For g(λ) = e^(−λ), g′(λ) = −e^(−λ). Therefore the Cramér–Rao lower bound is CRLB = [g′(λ)]²/I_n(λ) = e^(−2λ)/(n/λ) = λ e^(−2λ)/n.

Now let a = 1 − 1/n. Then E(W²) = E[a^(2T)] = exp[nλ(a² − 1)] = exp[nλ((1 − 1/n)² − 1)] = exp(−2λ + λ/n). So Var(W) = e^(−2λ+λ/n) − e^(−2λ) = e^(−2λ)(e^(λ/n) − 1). Since e^(λ/n) − 1 > λ/n for λ > 0, Var(W) > λ e^(−2λ)/n = CRLB. Thus W is UMVUE but does not attain the CRLB.

(b)(i) f_Y(y) = ∫_−∞^∞ f(x,y) dx = y^(3/2)e^(−y)/√(2π) ∫_−∞^∞ e^(−yx²/2) dx. Since ∫_−∞^∞ e^(−yx²/2) dx = √(2π/y), f_Y(y) = y e^(−y), y > 0. Thus Y ~ Gamma(shape = 2, rate = 1).

The conditional density is f_X|Y(x|y) = f(x,y)/f_Y(y) = √(y)/√(2π) e^(−yx²/2), so X|Y = y ~ N(0, 1/y).

(ii) For Y ~ Gamma(2,1): E(Y) = 2, V(Y) = 2. For X|Y = y ~ N(0, 1/y): E(X|Y) = 0, V(X|Y) = 1/Y. Equivalently, conditional on Y = y, V(X|Y=y) = 1/y.

(iii) By the law of total expectation, E(X) = E[E(X|Y)] = E(0) = 0. By the law of total variance, V(X) = E[V(X|Y)] + V[E(X|Y)] = E(1/Y) + 0. For Y ~ Gamma(2,1), E(1/Y) = ∫₀^∞ (1/y) y e^(−y) dy = ∫₀^∞ e^(−y) dy = 1. Therefore V(X) = 1.

(c) Use the Kruskal–Wallis test, since three independent samples are compared for equality of effectiveness. H₀: the three methods are equally effective. H₁: not all methods are equally effective.

The 18 scores in increasing order are: 60(B), 65(C), 67(B), 68(A), 72(B), 77(C), 78(C), 79(A), 80(A), 82(B), 83(A), 84(B), 85(A), 86(B), 88(C), 90(A), 91(B), 93(C).

Ranks by method: Method A: 4, 8, 9, 11, 13, 16; R₁ = 61. Method B: 1, 3, 5, 10, 12, 14, 17; R₂ = 62. Method C: 2, 6, 7, 15, 18; R₃ = 48.

Here N = 18, k = 3, n₁ = 6, n₂ = 7, n₃ = 5. The Kruskal–Wallis statistic is H = [12/(N(N+1))] Σⱼ Rⱼ²/nⱼ − 3(N+1) = (12/342)[61²/6 + 62²/7 + 48²/5] − 57 = 0.1968.

Under H₀, H is approximately χ² with k − 1 = 2 degrees of freedom. At 5% level, χ²₀.₀₅,₂ = 5.991. Since 0.1968 < 5.991, we fail to reject H₀.

Conclusion: At the 0.05 level of significance, there is no sufficient evidence to conclude that the three methods are not equally effective.

What "Prove" is asking you to do

Establish that the statement holds for every case it claims, not for one representative case. The argument must be closed: each line follows from a definition, a hypothesis, or a named theorem you are entitled to use.

Structure that answers it

Given and to prove, restated → theorem or construction to be used, named → the argument line by line → conclusion stated as proved

Where marks are lost

Testing one example, which illustrates but proves nothing. On an if and only if claim, proving one direction and stopping forfeits that half outright, and degenerate cases — zero, the empty set, the equality case — have to be disposed of rather than assumed away.

All UPSC directive words, compared →

How this answer will be evaluated

Approach

Framework: Statistical Inference & Non-Parametric Testing. (a) derive: given > assumptions > stepwise derivation > result > check | (b) derive: given > assumptions > stepwise derivation > result > check | (c) derive: given > assumptions > stepwise derivation > result > check Full marks: Rigorous proofs, correct formulas, clear interpretation, no calculation errors.

Key points expected

  • Poisson completeness proof
  • CRLB non-attainment logic
  • Marginal/conditional distribution extraction
  • Law of total variance
  • Kruskal-Wallis H statistic

Evaluation rubric

Each sub-part is marked on its own, against the marks and word limit printed on the paper.

  1. (a) Prove T is complete sufficient, unbiased for e^-λ, and fails CRLB. 20 marks

    derive— given → assumptions → stepwise derivation → result → check

    Must cover

    • Factorization theorem for sufficiency
    • Completeness via Poisson sum distribution
    • E(W) = e^-λ calculation
    • CRLB formula vs Var(W) comparison

    Loses marks

    • Skipping completeness proof
    • Incorrect CRLB formula

    Earns more

    • Explicit likelihood function
    • Correct CRLB derivation steps
    • Clear distinction of UMVUE vs CRLB

    Extra mark

    • Mention of Lehmann-Scheffé theorem
  2. (b) Find marginal/conditional distributions and moments for X and Y.

    derive— given → assumptions → stepwise derivation → result → check

    Must cover

    • Marginal f(y) integration
    • Conditional f(x|y) derivation
    • E(Y), V(Y) calculation
    • Law of total variance application

    Loses marks

    • Missing marginal derivation
    • Incorrect conditional distribution

    Earns more

    • Correct identification of Gamma/Normal forms
    • Step-by-step moment calculations

    Extra mark

    • Explicit integration limits
  3. (c) Apply Kruskal-Wallis test to compare three methods. 15 marks

    derive— given → assumptions → stepwise derivation → result → check

    Must cover

    • Ranking of all 18 observations
    • Sum of ranks per group
    • H statistic calculation
    • Chi-square critical value comparison

    Loses marks

    • Using ANOVA instead of non-parametric
    • Incorrect ranking or sum of ranks

    Earns more

    • Clean ranking table
    • Correct degrees of freedom (k-1)

    Extra mark

    • Mention of tie correction if applicable

Practice this exact question

Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.

Evaluate my answer →

More from Statistics 2021 Paper I