Paper I — Q1
(a) A production unit manufacturing surgical masks is concerned about the quality of their masks. A random sample of n masks are…
A production unit manufacturing surgical masks is concerned about the quality of their masks. A random sample of n masks are inspected to estimate 'p', the probability of manufacturing a defective mask. How large a sample is required so that the estimate of p lies in the range p ± 0.1 with probability 0.95 ? 10 marks
An insurance company studies a sample of 150 policy-holders. There are three categories of policies : auto, home and medical. The following results are obtained about the policies held by the policy-holders : 30 have only home insurance
10 have only medical insurance
98 have auto insurance, but not all three types of insurance
27 have medical insurance, but not all three types of insurance
13 have auto and medical insurance Given that a policy-holder has medical insurance, calculate the probability that he has home insurance. 10 marks
Let X and Y be independent and identically distributed exponential random variables with mean λ > 0. Define Z = 1, & if X < Y 0, & if X ≥ Y Find E[X|Z = 1] + E[X|Z = 0]. 10 marks
Let X₁, X₂, ..., Xₙ be a random sample from f(x, θ) = (log(θ))/(θ - 1)θ^x; 0 < x < 1, θ > 1 Is there a function of θ, say g(θ), for which there exists an unbiased estimator whose variance attains the C-R lower bound ? If yes, find it. If not, show why not. 10 marks
Let f(x, θ) be the Cauchy pdf f(x, θ) = (θ)/(π) 1/(θ² + x²); -∞ < x < ∞, θ > 0 Show that this family does not have Monotone Likelihood Ratio (MLR).
If X is one observation from f(x, θ), show that |X| is sufficient for θ and hence the distribution of |X| does have an MLR. (5+5 marks)
हिंदी में प्रश्न पढ़ें
सर्जिकल मास्क बनाने वाली एक उत्पादन इकाई की मास्क की गुणवत्ता जाँचने में रुचि है । दोषपूर्ण मास्क बनाने की प्रायिकता, 'p', के आकलन के लिए n मास्क के एक यादृच्छिक प्रतिदर्श का निरीक्षण किया गया । प्रतिदर्श कितना बड़ा होना चाहिए ताकि 0.95 प्रायिकता के साथ p के आकलक का परिसर p ± 0.1 हो ? (10 अंक)
एक बीमा कंपनी ने 150 पॉलिसी-धारकों के प्रतिदर्श का अध्ययन किया । पॉलिसी के तीन वर्ग हैं : वाहन, गृह और चिकित्सा । पॉलिसी-धारकों द्वारा गृहीत पॉलिसियों के संबंध में निम्न परिणाम प्राप्त हुए : 30 के पास केवल गृह बीमा है
10 के पास केवल चिकित्सा बीमा है
98 के पास वाहन बीमा है, लेकिन सभी तीन प्रकार के बीमा नहीं हैं
27 के पास चिकित्सा बीमा है, लेकिन सभी तीन प्रकार के बीमा नहीं हैं
13 के पास वाहन और चिकित्सा बीमा है यदि यह दिया हुआ है कि पॉलिसी-धारक के पास चिकित्सा बीमा है, तो उसके पास गृह बीमा होने की प्रायिकता परिकलित कीजिए । (10 अंक)
माना X और Y स्वतंत्र और सर्वसम बंटित चर्याताकि यादृच्छिक चर हैं जिनका माध्य λ > 0 है । परिभाषित है Z = 1, & यदि X < Y 0, & यदि X ≥ Y ज्ञात कीजिए : E[X|Z=1] + E[X|Z=0] (10 अंक)
माना X₁, X₂, ..., Xₙ f(x, θ) = (log(θ))/(θ - 1)θ^x; 0 < x < 1, θ > 1 से लिया गया कोई यादृच्छिक प्रतिदर्श है । क्या θ का एक फलन, माना g(θ), के लिए कोई अनभिनत आकलक है जिसका प्रसरण सी.-आर. (C-R) निम्न परिबंध प्राप्त करता हो ? यदि हाँ, तो ज्ञात कीजिए । यदि नहीं, तो दर्शाइए क्यों नहीं । (10 अंक)
माना f(x, θ) = (θ)/(π)1/(θ² + x²); -∞ < x < ∞, θ > 0 कौशी का प्रायिकता घनत्व फलन है । दर्शाइए कि इस बंटन संवर्ग के लिए एकदिष्ट संभाविता अनुपात (एम.एल.आर.) नहीं है ।
यदि X, f(x, θ) से लिया गया एक प्रेक्षण है, तो दर्शाइए कि |X|, θ का पर्याप्त प्रतिदर्शज है और इसलिए |X| के बंटन के लिए एम.एल.आर. है । (5+5 अंक)
Model answer
Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.
(a) Let p be the true proportion of defective masks and p̂ the sample proportion. For large n, by the normal approximation, p̂ ≈ N(p, p(1-p)/n). We need P(p̂ lies in p ± 0.1) = 0.95. Thus P(|p̂ − p| ≤ 0.1) = 0.95, so 0.1 = 1.96 sqrt(p(1-p)/n). Since p is unknown, use the maximum possible value p(1-p) ≤ 1/4. Hence n ≥ (1.96)² (1/4) / (0.1)² = 3.8416 × 0.25 / 0.01 = 96.04. Therefore the required sample size is n = 97 masks. If a prior guess p₀ is available, replace 1/4 by p₀(1-p₀). The normal approximation is valid when np and n(1-p) are at least about 5.
(b) Let A = auto, H = home, M = medical. Define: a = only auto, x = auto and home only, y = auto and medical only, z = home and medical only, t = all three. Given: h = only home = 30, m = only medical = 10. From (iii): a + x + y = 98. From (iv): m + y + z = 27 ⇒ 10 + y + z = 27 ⇒ y + z = 17. From (v): y + t = 13.
The total number of policy-holders is 150, so a + x + y + z + t + h + m = 150. Using a + x + y = 98 and h + m = 40, 98 + z + t + 40 = 150 ⇒ z + t = 12. Now y + z = 17 and z + t = 12 give y − t = 5, so y = t + 5. But y + t = 13, hence (t + 5) + t = 13 ⇒ 2t = 8 ⇒ t = 4. Then y = 9 and z = 8.
Thus the number having medical insurance is m + y + z + t = 10 + 9 + 8 + 4 = 31. The number having both home and medical insurance is z + t = 8 + 4 = 12. Therefore, P(H | M) = 12/31. Answer: 12/31.
(c) Since X and Y are iid exponential with mean λ, f_X(x) = (1/λ) e^−x/λ, x > 0. Also P(Z = 1) = P(X < Y) = 1/2.
Now E[X | Z = 1] = E[X | X < Y] = [∫₀^∞ ∫_x^∞ x f_X(x) f_Y(y) dy dx] / P(X < Y). The inner integral is f_X(x) P(Y > x) = (1/λ) e^−x/λ e^−x/λ. So the numerator is (1/λ) ∫₀^∞ x e^−2x/λ dx = (1/λ)(λ²/4) = λ/4. Dividing by 1/2 gives E[X | Z = 1] = λ/2.
Similarly, E[X | Z = 0] = [∫₀^∞ ∫₀^x x f_X(x) f_Y(y) dy dx] / P(X ≥ Y). The inner integral is f_X(x) F_X(x) = (1/λ) e^−x/λ(1 − e^−x/λ). Thus the numerator is (1/λ)[∫₀^∞ x e^−x/λ dx − ∫₀^∞ x e^−2x/λ dx] = (1/λ)(λ² − λ²/4) = 3λ/4. Dividing by 1/2 gives E[X | Z = 0] = 3λ/2.
Therefore, E[X | Z = 1] + E[X | Z = 0] = λ/2 + 3λ/2 = 2λ. Answer: 2λ.
(d) The density is f(x, θ) = log(θ)/(θ − 1) · θ^x, 0 < x < 1, θ > 1. Write η = log θ. Then f(x, θ) = η/(e^η − 1) e^ηx = exp(ηx − A(η)), where A(η) = log((e^η − 1)/η). Thus this is a one-parameter exponential family with sufficient statistic x. For a sample of size n, the likelihood depends on the data through S = Σ Xᵢ, and the family is regular.
For one observation, E[X] = A′(η) = e^η/(e^η − 1) − 1/η = θ/(θ − 1) − 1/log θ. Also Var(X) = A″(η) = 1/(log θ)² − θ/(θ − 1)².
Take g(θ) = E_θ[X] = θ/(θ − 1) − 1/log θ. Then X̄ = (1/n) Σ Xᵢ is unbiased for g(θ). Its variance is Var(X̄) = Var(X)/n. The derivative is g′(θ) = A″(η)/θ. The Fisher information for n observations is I_n(θ) = n A″(η)/θ². Therefore the Cramér–Rao lower bound for g(θ) is (g′(θ))² / I_n(θ) = (A″(η)²/θ²) / (n A″(η)/θ²) = A″(η)/n = Var(X)/n = Var(X̄). Hence the bound is attained.
Yes. The function is g(θ) = θ/(θ − 1) − 1/log θ, and X̄ is the unbiased estimator attaining the C-R lower bound.
(e)(i) For θ₂ > θ₁, the likelihood ratio is R(x) = f(x, θ₂)/f(x, θ₁) = [θ₂/(π(θ₂² + x²))] / [θ₁/(π(θ₁² + x²))] = (θ₂/θ₁) · (θ₁² + x²)/(θ₂² + x²). Then d/dx log R(x) = 2x/(θ₁² + x²) − 2x/(θ₂² + x²) = 2x(θ₂² − θ₁²) / [(θ₁² + x²)(θ₂² + x²)]. For x > 0 this derivative is positive, while for x < 0 it is negative. Hence R(x) decreases on (−∞, 0) and increases on (0, ∞). It is not monotone over the whole real line. Therefore the Cauchy family does not have a Monotone Likelihood Ratio.
(e)(ii) Since f(x, θ) = θ/(π(θ² + x²)) = θ/(π(θ² + |x|²)), the density depends on x only through |x|. By the factorization theorem, f(x, θ) = g(|x|, θ) h(x), with h(x) = 1 and g(t, θ) = θ/(π(θ² + t²)). Hence |X| is sufficient for θ.
Let Y = |X|. For y > 0, f_Y(y, θ) = f_X(y, θ) + f_X(−y, θ) = 2θ/(π(θ² + y²)). For θ₂ > θ₁, the likelihood ratio is r(y) = f_Y(y, θ₂)/f_Y(y, θ₁) = (θ₂/θ₁) · (θ₁² + y²)/(θ₂² + y²). Then d/dy log r(y) = 2y(θ₂² − θ₁²) / [(θ₁² + y²)(θ₂² + y²)] > 0 for y > 0. Thus r(y) is nondecreasing in y. Therefore the distribution of |X| does have a Monotone Likelihood Ratio.
What "Solve" is asking you to do
Choose the method, then carry it through to a final answer. Identifying what kind of problem this is and why that method applies is the first thing marked; a correct figure arrived at invisibly earns almost nothing.
Structure that answers it
Given data and what is required → method chosen, with the reason it applies → set-up (equation, circuit, free body, trial balance) → working, step by step → answer with units and any condition of validity
Where marks are lost
Doing the middle steps mentally and writing only the result. In mathematics papers, a further loss comes from giving a decimal where the exact value in surds or fractions was wanted, or from skipping the justification a part explicitly asks for.
How this answer will be evaluated
Approach
Framework: Statistical Inference & Probability Theory. (a) calculate: given > formula > substitution > result with units > interpretation | (b) calculate: given > formula > substitution > result with units > interpretation | (c) calculate: given > formula > substitution > result with units > interpretation | (d) calculate: given > formula > substitution > result with units > interpretation | (e) explain: definition/context > points in order > small example > short close Full marks: Rigorous derivations, correct set theory, clear MLR proofs, no calculation errors.
Key points expected
- Sample size formula n = (z^2 * 0.25) / E^2
- Venn diagram logic for conditional probability
- Integration over regions for conditional expectation
- Fisher Information and C-R bound conditions
- Likelihood ratio monotonicity test
Evaluation rubric
Each sub-part is marked on its own, against the marks and word limit printed on the paper.
- (a) Determine minimum sample size n for a 95% confidence interval with margin of error 0.1. 10 marks
calculate— given → formula → substitution → result with units → interpretation
Must cover
- State normal approximation for binomial proportion
- Identify z-score for 0.95 probability (1.96)
- Use worst-case variance p(1-p) = 0.25
- Solve inequality for n
Loses marks
- Using p=0.5 without justification
- Forgetting to square the z-score
Earns more
- Explicit formula for margin of error
- Rounding up to next integer
Extra mark
- Mentioning continuity correction
- (b) Compute conditional probability P(Home | Medical) using set theory/Venn diagram. 10 marks
calculate— given → formula → substitution → result with units → interpretation
Must cover
- Define sets A, H, M for Auto, Home, Medical
- Determine |M| (total medical holders)
- Determine |H ∩ M| (intersection of Home and Medical)
- Apply P(H|M) = |H ∩ M| / |M|
Loses marks
- Confusing 'only' with 'at least'
- Incorrectly calculating the intersection
Earns more
- Drawing a Venn diagram
- Step-by-step set subtraction
Extra mark
- Verifying total count sums to 150
- (c) Find the sum of conditional expectations E[X|X<Y] + E[X|X≥Y]. 10 marks
calculate— given → formula → substitution → result with units → interpretation
Must cover
- Use joint pdf of independent exponentials
- Set up integrals for E[X|Z=1] and E[X|Z=0]
- Evaluate integrals over triangular regions
- Sum the two results
Loses marks
- Incorrect conditional density normalization
- Calculation errors in integration
Earns more
- Using symmetry arguments
- Correct limits of integration
Extra mark
- Alternative method using order statistics
- (d) Determine if an unbiased estimator exists attaining the C-R lower bound. 10 marks
calculate— given → formula → substitution → result with units → interpretation
Must cover
- Calculate Fisher Information I(θ)
- Check regularity conditions for C-R bound
- Find MLE or sufficient statistic
- Verify if variance equals 1/I(θ)
Loses marks
- Assuming regularity without checking
- Confusing MLE with UMVUE
Earns more
- Explicit derivation of score function
- Checking if estimator is a function of sufficient statistic
Extra mark
- Identifying the specific function g(θ)
- (e) Show lack of MLR for Cauchy, then show |X| is sufficient and has MLR.
explain— definition/context → points in order → small example → short close
Must cover
- Show likelihood ratio is not monotone in x
- Apply Factorization Theorem for sufficiency of |X|
- Derive pdf of |X|
- Show pdf of |X| has MLR property
Loses marks
- Failing to show the ratio is not monotone
- Incorrect transformation of variables
Earns more
- Clear algebraic manipulation of likelihood ratio
- Explicit statement of Factorization Theorem
Extra mark
- Graphical illustration of non-monotonicity
Practice this exact question
Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.
Evaluate my answer →More from Statistics 2021 Paper I
- Q1 (a) A production unit manufacturing surgical masks is concerned about the quality of thei…
- Q2 (a) Let Y₁, Y₂, Y₃, ... be independent and identical Poisson random variables with parame…
- Q3 (a) Let X and Y be two independent random variables following exponential distribution wi…
- Q4 (a) Let X₁, X₂, ..., Xₙ be a random sample from Poisson distribution with mean λ > 0. Def…