Statistics 2021 Paper I 50 marks Compulsory Solve

Paper I — Q1

(a) A production unit manufacturing surgical masks is concerned about the quality of their masks. A random sample of n masks are…

(a)

A production unit manufacturing surgical masks is concerned about the quality of their masks. A random sample of n masks are inspected to estimate 'p', the probability of manufacturing a defective mask. How large a sample is required so that the estimate of p lies in the range p ± 0.1 with probability 0.95 ? 10 marks

(b)
(i)

An insurance company studies a sample of 150 policy-holders. There are three categories of policies : auto, home and medical. The following results are obtained about the policies held by the policy-holders : 30 have only home insurance

(ii)

10 have only medical insurance

(iii)

98 have auto insurance, but not all three types of insurance

(iv)

27 have medical insurance, but not all three types of insurance

(v)

13 have auto and medical insurance Given that a policy-holder has medical insurance, calculate the probability that he has home insurance. 10 marks

(c)

Let X and Y be independent and identically distributed exponential random variables with mean λ > 0. Define Z = 1, & if X < Y 0, & if X ≥ Y Find E[X|Z = 1] + E[X|Z = 0]. 10 marks

(d)

Let X₁, X₂, ..., Xₙ be a random sample from f(x, θ) = (log(θ))/(θ - 1)θ^x; 0 < x < 1, θ > 1 Is there a function of θ, say g(θ), for which there exists an unbiased estimator whose variance attains the C-R lower bound ? If yes, find it. If not, show why not. 10 marks

(e)
(i)

Let f(x, θ) be the Cauchy pdf f(x, θ) = (θ)/(π) 1/(θ² + x²); -∞ < x < ∞, θ > 0 Show that this family does not have Monotone Likelihood Ratio (MLR).

(ii)

If X is one observation from f(x, θ), show that |X| is sufficient for θ and hence the distribution of |X| does have an MLR. (5+5 marks)

हिंदी में प्रश्न पढ़ें
(a)

सर्जिकल मास्क बनाने वाली एक उत्पादन इकाई की मास्क की गुणवत्ता जाँचने में रुचि है । दोषपूर्ण मास्क बनाने की प्रायिकता, 'p', के आकलन के लिए n मास्क के एक यादृच्छिक प्रतिदर्श का निरीक्षण किया गया । प्रतिदर्श कितना बड़ा होना चाहिए ताकि 0.95 प्रायिकता के साथ p के आकलक का परिसर p ± 0.1 हो ? (10 अंक)

(b)
(i)

एक बीमा कंपनी ने 150 पॉलिसी-धारकों के प्रतिदर्श का अध्ययन किया । पॉलिसी के तीन वर्ग हैं : वाहन, गृह और चिकित्सा । पॉलिसी-धारकों द्वारा गृहीत पॉलिसियों के संबंध में निम्न परिणाम प्राप्त हुए : 30 के पास केवल गृह बीमा है

(ii)

10 के पास केवल चिकित्सा बीमा है

(iii)

98 के पास वाहन बीमा है, लेकिन सभी तीन प्रकार के बीमा नहीं हैं

(iv)

27 के पास चिकित्सा बीमा है, लेकिन सभी तीन प्रकार के बीमा नहीं हैं

(v)

13 के पास वाहन और चिकित्सा बीमा है यदि यह दिया हुआ है कि पॉलिसी-धारक के पास चिकित्सा बीमा है, तो उसके पास गृह बीमा होने की प्रायिकता परिकलित कीजिए । (10 अंक)

(c)

माना X और Y स्वतंत्र और सर्वसम बंटित चर्याताकि यादृच्छिक चर हैं जिनका माध्य λ > 0 है । परिभाषित है Z = 1, & यदि X < Y 0, & यदि X ≥ Y ज्ञात कीजिए : E[X|Z=1] + E[X|Z=0] (10 अंक)

(d)

माना X₁, X₂, ..., Xₙ f(x, θ) = (log(θ))/(θ - 1)θ^x; 0 < x < 1, θ > 1 से लिया गया कोई यादृच्छिक प्रतिदर्श है । क्या θ का एक फलन, माना g(θ), के लिए कोई अनभिनत आकलक है जिसका प्रसरण सी.-आर. (C-R) निम्न परिबंध प्राप्त करता हो ? यदि हाँ, तो ज्ञात कीजिए । यदि नहीं, तो दर्शाइए क्यों नहीं । (10 अंक)

(e)
(i)

माना f(x, θ) = (θ)/(π)1/(θ² + x²); -∞ < x < ∞, θ > 0 कौशी का प्रायिकता घनत्व फलन है । दर्शाइए कि इस बंटन संवर्ग के लिए एकदिष्ट संभाविता अनुपात (एम.एल.आर.) नहीं है ।

(ii)

यदि X, f(x, θ) से लिया गया एक प्रेक्षण है, तो दर्शाइए कि |X|, θ का पर्याप्त प्रतिदर्शज है और इसलिए |X| के बंटन के लिए एम.एल.आर. है । (5+5 अंक)

Q1 of the 2021 UPSC Mains Statistics Paper I, as printed
The question as printed in the 2021 Statistics paper

Model answer

Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.

(a) Let p be the true proportion of defective masks and p̂ the sample proportion. For large n, by the normal approximation, p̂ ≈ N(p, p(1-p)/n). We need P(p̂ lies in p ± 0.1) = 0.95. Thus P(|p̂ − p| ≤ 0.1) = 0.95, so 0.1 = 1.96 sqrt(p(1-p)/n). Since p is unknown, use the maximum possible value p(1-p) ≤ 1/4. Hence n ≥ (1.96)² (1/4) / (0.1)² = 3.8416 × 0.25 / 0.01 = 96.04. Therefore the required sample size is n = 97 masks. If a prior guess p₀ is available, replace 1/4 by p₀(1-p₀). The normal approximation is valid when np and n(1-p) are at least about 5.

(b) Let A = auto, H = home, M = medical. Define: a = only auto, x = auto and home only, y = auto and medical only, z = home and medical only, t = all three. Given: h = only home = 30, m = only medical = 10. From (iii): a + x + y = 98. From (iv): m + y + z = 27 ⇒ 10 + y + z = 27 ⇒ y + z = 17. From (v): y + t = 13.

The total number of policy-holders is 150, so a + x + y + z + t + h + m = 150. Using a + x + y = 98 and h + m = 40, 98 + z + t + 40 = 150 ⇒ z + t = 12. Now y + z = 17 and z + t = 12 give y − t = 5, so y = t + 5. But y + t = 13, hence (t + 5) + t = 13 ⇒ 2t = 8 ⇒ t = 4. Then y = 9 and z = 8.

Thus the number having medical insurance is m + y + z + t = 10 + 9 + 8 + 4 = 31. The number having both home and medical insurance is z + t = 8 + 4 = 12. Therefore, P(H | M) = 12/31. Answer: 12/31.

(c) Since X and Y are iid exponential with mean λ, f_X(x) = (1/λ) e^−x/λ, x > 0. Also P(Z = 1) = P(X < Y) = 1/2.

Now E[X | Z = 1] = E[X | X < Y] = [∫₀^∞ ∫_x^∞ x f_X(x) f_Y(y) dy dx] / P(X < Y). The inner integral is f_X(x) P(Y > x) = (1/λ) e^−x/λ e^−x/λ. So the numerator is (1/λ) ∫₀^∞ x e^−2x/λ dx = (1/λ)(λ²/4) = λ/4. Dividing by 1/2 gives E[X | Z = 1] = λ/2.

Similarly, E[X | Z = 0] = [∫₀^∞ ∫₀^x x f_X(x) f_Y(y) dy dx] / P(X ≥ Y). The inner integral is f_X(x) F_X(x) = (1/λ) e^−x/λ(1 − e^−x/λ). Thus the numerator is (1/λ)[∫₀^∞ x e^−x/λ dx − ∫₀^∞ x e^−2x/λ dx] = (1/λ)(λ² − λ²/4) = 3λ/4. Dividing by 1/2 gives E[X | Z = 0] = 3λ/2.

Therefore, E[X | Z = 1] + E[X | Z = 0] = λ/2 + 3λ/2 = 2λ. Answer: 2λ.

(d) The density is f(x, θ) = log(θ)/(θ − 1) · θ^x, 0 < x < 1, θ > 1. Write η = log θ. Then f(x, θ) = η/(e^η − 1) e^ηx = exp(ηx − A(η)), where A(η) = log((e^η − 1)/η). Thus this is a one-parameter exponential family with sufficient statistic x. For a sample of size n, the likelihood depends on the data through S = Σ Xᵢ, and the family is regular.

For one observation, E[X] = A′(η) = e^η/(e^η − 1) − 1/η = θ/(θ − 1) − 1/log θ. Also Var(X) = A″(η) = 1/(log θ)² − θ/(θ − 1)².

Take g(θ) = E_θ[X] = θ/(θ − 1) − 1/log θ. Then X̄ = (1/n) Σ Xᵢ is unbiased for g(θ). Its variance is Var(X̄) = Var(X)/n. The derivative is g′(θ) = A″(η)/θ. The Fisher information for n observations is I_n(θ) = n A″(η)/θ². Therefore the Cramér–Rao lower bound for g(θ) is (g′(θ))² / I_n(θ) = (A″(η)²/θ²) / (n A″(η)/θ²) = A″(η)/n = Var(X)/n = Var(X̄). Hence the bound is attained.

Yes. The function is g(θ) = θ/(θ − 1) − 1/log θ, and X̄ is the unbiased estimator attaining the C-R lower bound.

(e)(i) For θ₂ > θ₁, the likelihood ratio is R(x) = f(x, θ₂)/f(x, θ₁) = [θ₂/(π(θ₂² + x²))] / [θ₁/(π(θ₁² + x²))] = (θ₂/θ₁) · (θ₁² + x²)/(θ₂² + x²). Then d/dx log R(x) = 2x/(θ₁² + x²) − 2x/(θ₂² + x²) = 2x(θ₂² − θ₁²) / [(θ₁² + x²)(θ₂² + x²)]. For x > 0 this derivative is positive, while for x < 0 it is negative. Hence R(x) decreases on (−∞, 0) and increases on (0, ∞). It is not monotone over the whole real line. Therefore the Cauchy family does not have a Monotone Likelihood Ratio.

(e)(ii) Since f(x, θ) = θ/(π(θ² + x²)) = θ/(π(θ² + |x|²)), the density depends on x only through |x|. By the factorization theorem, f(x, θ) = g(|x|, θ) h(x), with h(x) = 1 and g(t, θ) = θ/(π(θ² + t²)). Hence |X| is sufficient for θ.

Let Y = |X|. For y > 0, f_Y(y, θ) = f_X(y, θ) + f_X(−y, θ) = 2θ/(π(θ² + y²)). For θ₂ > θ₁, the likelihood ratio is r(y) = f_Y(y, θ₂)/f_Y(y, θ₁) = (θ₂/θ₁) · (θ₁² + y²)/(θ₂² + y²). Then d/dy log r(y) = 2y(θ₂² − θ₁²) / [(θ₁² + y²)(θ₂² + y²)] > 0 for y > 0. Thus r(y) is nondecreasing in y. Therefore the distribution of |X| does have a Monotone Likelihood Ratio.

What "Solve" is asking you to do

Choose the method, then carry it through to a final answer. Identifying what kind of problem this is and why that method applies is the first thing marked; a correct figure arrived at invisibly earns almost nothing.

Structure that answers it

Given data and what is required → method chosen, with the reason it applies → set-up (equation, circuit, free body, trial balance) → working, step by step → answer with units and any condition of validity

Where marks are lost

Doing the middle steps mentally and writing only the result. In mathematics papers, a further loss comes from giving a decimal where the exact value in surds or fractions was wanted, or from skipping the justification a part explicitly asks for.

All UPSC directive words, compared →

How this answer will be evaluated

Approach

Framework: Statistical Inference & Probability Theory. (a) calculate: given > formula > substitution > result with units > interpretation | (b) calculate: given > formula > substitution > result with units > interpretation | (c) calculate: given > formula > substitution > result with units > interpretation | (d) calculate: given > formula > substitution > result with units > interpretation | (e) explain: definition/context > points in order > small example > short close Full marks: Rigorous derivations, correct set theory, clear MLR proofs, no calculation errors.

Key points expected

  • Sample size formula n = (z^2 * 0.25) / E^2
  • Venn diagram logic for conditional probability
  • Integration over regions for conditional expectation
  • Fisher Information and C-R bound conditions
  • Likelihood ratio monotonicity test

Evaluation rubric

Each sub-part is marked on its own, against the marks and word limit printed on the paper.

  1. (a) Determine minimum sample size n for a 95% confidence interval with margin of error 0.1. 10 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • State normal approximation for binomial proportion
    • Identify z-score for 0.95 probability (1.96)
    • Use worst-case variance p(1-p) = 0.25
    • Solve inequality for n

    Loses marks

    • Using p=0.5 without justification
    • Forgetting to square the z-score

    Earns more

    • Explicit formula for margin of error
    • Rounding up to next integer

    Extra mark

    • Mentioning continuity correction
  2. (b) Compute conditional probability P(Home | Medical) using set theory/Venn diagram. 10 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • Define sets A, H, M for Auto, Home, Medical
    • Determine |M| (total medical holders)
    • Determine |H ∩ M| (intersection of Home and Medical)
    • Apply P(H|M) = |H ∩ M| / |M|

    Loses marks

    • Confusing 'only' with 'at least'
    • Incorrectly calculating the intersection

    Earns more

    • Drawing a Venn diagram
    • Step-by-step set subtraction

    Extra mark

    • Verifying total count sums to 150
  3. (c) Find the sum of conditional expectations E[X|X<Y] + E[X|X≥Y]. 10 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • Use joint pdf of independent exponentials
    • Set up integrals for E[X|Z=1] and E[X|Z=0]
    • Evaluate integrals over triangular regions
    • Sum the two results

    Loses marks

    • Incorrect conditional density normalization
    • Calculation errors in integration

    Earns more

    • Using symmetry arguments
    • Correct limits of integration

    Extra mark

    • Alternative method using order statistics
  4. (d) Determine if an unbiased estimator exists attaining the C-R lower bound. 10 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • Calculate Fisher Information I(θ)
    • Check regularity conditions for C-R bound
    • Find MLE or sufficient statistic
    • Verify if variance equals 1/I(θ)

    Loses marks

    • Assuming regularity without checking
    • Confusing MLE with UMVUE

    Earns more

    • Explicit derivation of score function
    • Checking if estimator is a function of sufficient statistic

    Extra mark

    • Identifying the specific function g(θ)
  5. (e) Show lack of MLR for Cauchy, then show |X| is sufficient and has MLR.

    explain— definition/context → points in order → small example → short close

    Must cover

    • Show likelihood ratio is not monotone in x
    • Apply Factorization Theorem for sufficiency of |X|
    • Derive pdf of |X|
    • Show pdf of |X| has MLR property

    Loses marks

    • Failing to show the ratio is not monotone
    • Incorrect transformation of variables

    Earns more

    • Clear algebraic manipulation of likelihood ratio
    • Explicit statement of Factorization Theorem

    Extra mark

    • Graphical illustration of non-monotonicity

Practice this exact question

Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.

Evaluate my answer →

More from Statistics 2021 Paper I