Statistics

UPSC Statistics 2021 — Paper I

All 8 questions from UPSC Civil Services Mains Statistics 2021 Paper I (400 marks total). Every stem reproduced in full, with directive-word analysis, marks, word limits, and answer-approach pointers.

8Questions
400Total marks
2021Year
Paper IPaper

Topics covered

Probability theory and statistical inference (1)Limit theorems and characteristic functions (1)Sequential probability ratio test and convergence (1)UMVUE, joint distributions and non-parametric tests (1)Linear regression, experimental designs, sampling theory (1)Multivariate analysis, correlation, cluster sampling, multivariate normal distribution (1)Factorial experiments, principal components, regression estimator (1)Stratified sampling, polynomial regression, split-plot designs (1)

A

Q1
50M Compulsory solve Probability theory and statistical inference

(a) A production unit manufacturing surgical masks is concerned about the quality of their masks. A random sample of n masks are inspected to estimate 'p', the probability of manufacturing a defective mask. How large a sample is required so that the estimate of p lies in the range p ± 0.1 with probability 0.95 ? (10 marks) (b) An insurance company studies a sample of 150 policy-holders. There are three categories of policies : auto, home and medical. The following results are obtained about the policies held by the policy-holders : (i) 30 have only home insurance (ii) 10 have only medical insurance (iii) 98 have auto insurance, but not all three types of insurance (iv) 27 have medical insurance, but not all three types of insurance (v) 13 have auto and medical insurance Given that a policy-holder has medical insurance, calculate the probability that he has home insurance. (10 marks) (c) Let X and Y be independent and identically distributed exponential random variables with mean λ > 0. Define Z = 1, & if X < Y 0, & if X ≥ Y Find E[X|Z = 1] + E[X|Z = 0]. (10 marks) (d) Let X₁, X₂, ..., Xₙ be a random sample from f(x, θ) = (log(θ))/(θ - 1)θ^x; 0 < x < 1, θ > 1 Is there a function of θ, say g(θ), for which there exists an unbiased estimator whose variance attains the C-R lower bound ? If yes, find it. If not, show why not. (10 marks) (e) Let f(x, θ) be the Cauchy pdf f(x, θ) = (θ)/(π) 1/(θ² + x²); -∞ < x < ∞, θ > 0 (i) Show that this family does not have Monotone Likelihood Ratio (MLR). (ii) If X is one observation from f(x, θ), show that |X| is sufficient for θ and hence the distribution of |X| does have an MLR. (5+5 marks)

हिंदी में पढ़ें

(a) सर्जिकल मास्क बनाने वाली एक उत्पादन इकाई की मास्क की गुणवत्ता जाँचने में रुचि है । दोषपूर्ण मास्क बनाने की प्रायिकता, 'p', के आकलन के लिए n मास्क के एक यादृच्छिक प्रतिदर्श का निरीक्षण किया गया । प्रतिदर्श कितना बड़ा होना चाहिए ताकि 0.95 प्रायिकता के साथ p के आकलक का परिसर p ± 0.1 हो ? (10 अंक) (b) एक बीमा कंपनी ने 150 पॉलिसी-धारकों के प्रतिदर्श का अध्ययन किया । पॉलिसी के तीन वर्ग हैं : वाहन, गृह और चिकित्सा । पॉलिसी-धारकों द्वारा गृहीत पॉलिसियों के संबंध में निम्न परिणाम प्राप्त हुए : (i) 30 के पास केवल गृह बीमा है (ii) 10 के पास केवल चिकित्सा बीमा है (iii) 98 के पास वाहन बीमा है, लेकिन सभी तीन प्रकार के बीमा नहीं हैं (iv) 27 के पास चिकित्सा बीमा है, लेकिन सभी तीन प्रकार के बीमा नहीं हैं (v) 13 के पास वाहन और चिकित्सा बीमा है यदि यह दिया हुआ है कि पॉलिसी-धारक के पास चिकित्सा बीमा है, तो उसके पास गृह बीमा होने की प्रायिकता परिकलित कीजिए । (10 अंक) (c) माना X और Y स्वतंत्र और सर्वसम बंटित चर्याताकि यादृच्छिक चर हैं जिनका माध्य λ > 0 है । परिभाषित है Z = 1, & यदि X < Y 0, & यदि X ≥ Y ज्ञात कीजिए : E[X|Z=1] + E[X|Z=0] (10 अंक) (d) माना X₁, X₂, ..., Xₙ f(x, θ) = (log(θ))/(θ - 1)θ^x; 0 < x < 1, θ > 1 से लिया गया कोई यादृच्छिक प्रतिदर्श है । क्या θ का एक फलन, माना g(θ), के लिए कोई अनभिनत आकलक है जिसका प्रसरण सी.-आर. (C-R) निम्न परिबंध प्राप्त करता हो ? यदि हाँ, तो ज्ञात कीजिए । यदि नहीं, तो दर्शाइए क्यों नहीं । (10 अंक) (e) माना f(x, θ) = (θ)/(π)1/(θ² + x²); -∞ < x < ∞, θ > 0 कौशी का प्रायिकता घनत्व फलन है । (i) दर्शाइए कि इस बंटन संवर्ग के लिए एकदिष्ट संभाविता अनुपात (एम.एल.आर.) नहीं है । (ii) यदि X, f(x, θ) से लिया गया एक प्रेक्षण है, तो दर्शाइए कि |X|, θ का पर्याप्त प्रतिदर्शज है और इसलिए |X| के बंटन के लिए एम.एल.आर. है । (5+5 अंक)

Answer approach & key points

Framework: Statistical Inference & Probability Theory. (a) calculate: given > formula > substitution > result with units > interpretation | (b) calculate: given > formula > substitution > result with units > interpretation | (c) calculate: given > formula > substitution > result with units > interpretation | (d) calculate: given > formula > substitution > result with units > interpretation | (e) explain: definition/context > points in order > small example > short close Full marks: Rigorous derivations, correct set theory, clear MLR proofs, no calculation errors.

  • Sample size formula n = (z^2 * 0.25) / E^2
  • Venn diagram logic for conditional probability
  • Integration over regions for conditional expectation
  • Fisher Information and C-R bound conditions
  • Likelihood ratio monotonicity test
Q2
50M derive Limit theorems and characteristic functions

(a) Let Y₁, Y₂, Y₃, ... be independent and identical Poisson random variables with parameter 1. Use central limit theorem to establish n! ≃ √(2π n)(n/e)ⁿ for large value of positive integer n. (20 marks) (b) Let X₁, X₂, ..., Xₙ be a random sample such that log Xᵢ ~ N(θ, θ) distribution with θ > 0 unknown. Show that one of the solutions of the likelihood equation is the unique MLE of θ. Obtain asymptotic distribution of MLE of θ. (15 marks) (c) (i) State the sufficient conditions for a function φ(t) to be a characteristic function. (ii) Investigate if the following functions are characteristic functions : 1. e⁻ᵗ⁴ 2. [1 + |t|]⁻¹ Justify your answer. (5+10 marks)

हिंदी में पढ़ें

(a) माना Y₁, Y₂, Y₃, ... स्वतंत्र और सर्वसम व्यासों यादृच्छिक चर हैं जिनका प्राचल 1 है। केन्द्रीय सीमा प्रमेय का उपयोग करते हुए स्थापित कीजिए n! ≃ √(2π n)(n/e)ⁿ, जबकि धनात्मक पूर्णांक n बहुत है। (20 अंक) (b) माना X₁, X₂, ..., Xₙ ऐसा यादृच्छिक प्रतिदर्श है जिसका बंटन log Xᵢ ~ N(θ, θ), θ > 0 अज्ञात है। दर्शाइए कि संभाविता समीकरण का एक हल θ का एकमात्र MLE है। θ के MLE के लिए उपगामी बंटन प्राप्त कीजिए। (15 अंक) (c) (i) फलन φ(t) के अभिलक्षण फलन होने के लिए पर्याप्त प्रतिबंधों को लिखिए। (ii) जाँच कीजिए कि क्या निम्न फलन अभिलक्षण फलन हैं : 1. e⁻ᵗ⁴ 2. [1 + |t|]⁻¹ अपने उत्तर को तर्कसंगत सिद्ध कीजिए। (5+10 अंक)

Answer approach & key points

Framework: null. (a) derive: given > assumptions > stepwise derivation > result > check | (b) derive: given > assumptions > stepwise derivation > result > check | (c(i)) enumerate: list the items in order > one line each > no commentary | (c(ii)) justify: claim > 3-4 reasons > evidence > conclusion Full marks: Rigorous derivation with all steps shown and correct asymptotic results.

  • Define S_n = sum of n i.i.d. Poisson(1) variables
  • State S_n ~ Poisson(n) and apply CLT to standardize
  • Relate P(S_n = n) to the normal density at the mean
  • Equate the Poisson probability n! term with the normal density
  • Write the log-likelihood function for log X_i ~ N(theta, theta)
  • Differentiate to find the likelihood equation
  • Identify the unique solution as the MLE
  • State the asymptotic normal distribution using Fisher information
Q3
50M construct Sequential probability ratio test and convergence

(a) Let X and Y be two independent random variables following exponential distribution with mean 1/(λ) and 1/(μ) respectively, λ > 0, μ > 0. Suppose that (X₁, X₂, ..., Xₙ) and (Y₁, Y₂, ..., Yₙ) are sequences of observations on X and Y respectively. A random variable Uᵢ is defined as Uᵢ = 1, & if Xᵢ ≥ Yᵢ, i = 1, 2, ..., n 0, & otherwise Construct Wald's SPRT procedure based on Uᵢ's for testing H : λ = μ versus K : λ = 2μ with strength (α, β). (20 marks) (b) Let Yᵢ, i ≥ 1 be independent and identical U(-1, 1) random variables. Determine if the following sequences converge in probability : (i) (Yᵢ)/i (ii) (Yᵢ)ⁱ (5+10 marks) (c) Let X₁, X₂, ..., Xₙ be a random sample from uniform distribution U(− θ, θ), θ > 0. Find the complete sufficient statistic for θ. Hence, obtain the best unbiased estimator of θ. (15 marks)

हिंदी में पढ़ें

(a) माना X और Y चर्यातांकी बंटन से लिए गए दो स्वतंत्र यादृच्छिक चर हैं जिनका माध्य क्रमशः 1/(λ) और 1/(μ), λ > 0, μ > 0 है । माना (X₁, X₂, ..., Xₙ) और (Y₁, Y₂, ..., Yₙ) क्रमशः X और Y से लिए गए प्रेक्षणों के अनुक्रम हैं । एक यादृच्छिक चर Uᵢ इस प्रकार से परिभाषित है Uᵢ = 1, & यदि Xᵢ ≥ Yᵢ, i = 1, 2, ..., n 0, & अन्यथा Uᵢ पर आधारित H : λ = μ विरुद्ध K : λ = 2μ के परीक्षण के लिए वाल्ड SPRT विधि की रचना कीजिए जिसकी शक्ति (α, β) है । (20 अंक) (b) माना Yᵢ, i ≥ 1, स्वतंत्र और सर्वसम U(-1, 1) यादृच्छिक चर हैं । ज्ञात कीजिए कि क्या निम्न अनुक्रम प्रायिकता में अभिसरित हैं : (i) (Yᵢ)/i (ii) (Yᵢ)ⁱ (5+10 अंक) (c) माना X₁, X₂, ..., Xₙ एकसमान बंटन U(− θ, θ), θ > 0 से लिया गया एक यादृच्छिक प्रतिदर्श है । θ का पूर्ण पर्याप्त प्रतिदर्शज्ञात कीजिए । इससे θ का सर्वोत्तम अनभिनत आकलक प्राप्त कीजिए । (15 अंक)

Answer approach & key points

Framework: Wald's Sequential Probability Ratio Test (SPRT). (a) calculate: given > formula > substitution > result with units > interpretation | (b) calculate: given > formula > substitution > result with units > interpretation | (c) calculate: given > formula > substitution > result with units > interpretation Full marks: Rigorous derivation with all steps shown and correct interpretation.

  • P(U_i=1) = lambda/(lambda+mu)
  • Likelihood ratio L_n = (lambda/(lambda+mu))^S * (mu/(lambda+mu))^(n-S)
  • Convergence in probability definition
  • M = max(|X_i|) is complete sufficient
  • UMVUE = (n+1)/n * M
Q4
50M prove UMVUE, joint distributions and non-parametric tests

(a) Let X₁, X₂, ..., Xₙ be a random sample from Poisson distribution with mean λ > 0. Define a statistic W = (1 − 1/n)^T, T = Σᵢ₌₁ⁿ Xᵢ (i) Show that T is complete sufficient statistic. (ii) Show that T is unbiased for e^(−λ). (iii) Show that even though T is UMVUE, it does not attain the CRLB for g(λ) = e^(−λ). (20 marks) (b) Let f(x, y) = (e^(-yx²)/2 y³/2 e^-y)/(√(2π)), -∞ < x < ∞, y > 0. (i) Obtain the marginal distribution of Y and conditional distribution of X given Y. (ii) Find E(Y), V(Y), E(X|Y), V(X|Y). (iii) Use (ii) to find E(X), V(X). (5+5+5 marks) (c) A company's trainees are randomly assigned to groups which are through a certain industrial inspection procedure by three different methods. At the end of the instructing period they are tested for inspection performance quality. The following are their scores : Method A : 80 83 79 85 90 68 Method B : 82 84 60 72 86 67 91 Method C : 93 65 77 78 88 Using the appropriate non-parametric test, determine at 0·05 level of significance whether the three methods are equally effective. (15 marks)

हिंदी में पढ़ें

(a) माना X₁, X₂, ..., Xₙ प्वासों बंटन, जिसका माध्य λ > 0, से लिया गया एक यादृच्छिक प्रतिदर्श है । एक प्रतिदर्शज परिभाषित है W = (1 − 1/n)^T, T = Σᵢ₌₁ⁿ Xᵢ (i) दर्शाइए कि T पूर्ण पर्याप्त प्रतिदर्शज है । (ii) दर्शाइए कि T, e^(−λ) के लिए अनभिनत है । (iii) भले ही T, UMVUE (यू.एम.वी.यू.ई.) है, दर्शाइए कि यह g(λ) = e^(−λ) के लिए CRLB (सी.आर.एल.बी.) प्राप्त नहीं करता है । (20 अंक) (b) माना f(x, y) = [e^(−yx²/2) y^(3/2) e^(−y)] / √(2π) , −∞ < x < ∞, y > 0. (i) Y का उपांत बंटन और Y के दिए होने पर X का सप्रतिबंध बंटन प्राप्त कीजिए । (ii) E(Y), V(Y), E(X|Y), V(X|Y) ज्ञात कीजिए । (iii) (ii) का उपयोग करते हुए E(X), V(X) ज्ञात कीजिए । (5+5+5 अंक) (c) तीन विभिन्न तरीकों से एक निश्चित औद्योगिक निरीक्षण प्रक्रिया द्वारा एक कंपनी के प्रशिक्षार्थियों को यादृच्छया समूहों में नियत किया गया । प्रशिक्षण अवधि की समाप्ति पर निरीक्षण प्रदर्शन गुणवत्ता के लिए उनका परीक्षण किया गया । उनके स्कोर निम्न हैं : रीति A : 80 83 79 85 90 68 रीति B : 82 84 60 72 86 67 91 रीति C : 93 65 77 78 88 उपयुक्त अप्राचलिक परीक्षण का उपयोग करते हुए, 0·05 सार्थकता स्तर पर निर्धारित कीजिए कि क्या तीनों रीतियाँ समान रूप से प्रभावी हैं । (15 अंक)

Answer approach & key points

Framework: Statistical Inference & Non-Parametric Testing. (a) derive: given > assumptions > stepwise derivation > result > check | (b) derive: given > assumptions > stepwise derivation > result > check | (c) derive: given > assumptions > stepwise derivation > result > check Full marks: Rigorous proofs, correct formulas, clear interpretation, no calculation errors.

  • Poisson completeness proof
  • CRLB non-attainment logic
  • Marginal/conditional distribution extraction
  • Law of total variance
  • Kruskal-Wallis H statistic

B

Q5
50M Compulsory derive Linear regression, experimental designs, sampling theory

(a) For a simple linear regression model Y = β₀ + β₁Xᵢ + εᵢ, i = 1, ..., n (i) Derive the least square estimators of β₀ and β₁, clearly stating the conditions assumed. (ii) For eᵢ = Yᵢ - Ŷᵢ where Ŷᵢ is the fitted value, show that 1. Σᵢ₌₁ⁿ eᵢ = 0 2. Σᵢ₌₁ⁿ Yᵢ = Σᵢ₌₁ⁿ Ŷᵢ 3. Σᵢ₌₁ⁿ Xᵢeᵢ = 0 4. Σᵢ₌₁ⁿ Ŷᵢeᵢ = 0 5. The regression line passes through (X̄, Ȳ). 5+5 (b) In usual notations, if v, b, r, k and λ are the parameters of a Balanced Incomplete Block Design, then show that : (i) b ≥ r + 1 ≥ λ + 2 (ii) v ≤ b ≤ (r² - 1)/λ 10 (c) For the multiple linear regression model with two predictor variables X₁ and X₂, show that the estimate of regression coefficient of X₁ is unchanged when X₂ is added to the regression model, whenever X₁ and X₂ are uncorrelated. 10 (d) A sample of size n is drawn from a population having N units by simple random sampling without replacement. A sub-sample of n₁ units is drawn from the n units by simple random sampling without replacement. Let ȳ₁ denote the mean based on n₁ units and ȳ₂, the mean based on n₂ = n - n₁ units. Consider the estimator of the population mean Ȳₙ given by : Ŷₙ = wȳ₁ + (1-w)ȳ₂ ; 0 < w < 1 Show that E(Ŷₙ) = Ȳₙ, and obtain its variance. 10 (e) How is the efficiency of a design measured ? Derive the expression to measure the efficiency of a Randomised Block Design over a Completely Randomised Design. 10

हिंदी में पढ़ें

(a) एक साधारण रैखिक समाश्रयण निदर्श Y = β₀ + β₁Xᵢ + εᵢ, i = 1, ..., n के लिए (i) माने गए प्रतिबंधों को स्पष्ट लिखते हुए, β₀ और β₁ के न्यूनतम वर्ग आकलकों को व्युत्पन्न कीजिए। (ii) eᵢ = Yᵢ - Ŷᵢ जहाँ Ŷᵢ आसंजित मान है, के लिए दर्शाइए कि 1. Σᵢ₌₁ⁿ eᵢ = 0 2. Σᵢ₌₁ⁿ Yᵢ = Σᵢ₌₁ⁿ Ŷᵢ 3. Σᵢ₌₁ⁿ Xᵢeᵢ = 0 4. Σᵢ₌₁ⁿ Ŷᵢeᵢ = 0 5. समाश्रयण रेखा (X̄, Ȳ) से गुजरती है। 5+5 (b) प्रचलित संकेतों में, यदि v, b, r, k और λ किसी संतुलित अपूर्ण खंडक अभिकल्पना के प्राचल हैं, तो दर्शाइए कि : (i) b ≥ r + 1 ≥ λ + 2 (ii) v ≤ b ≤ (r² - 1)/λ 10 (c) एक बहु रैखिक समाश्रयण निदर्श जिसमें X₁ और X₂ दो प्रावकता चर हैं, के लिए दर्शाइए कि जब भी X₁ और X₂ असहसंबंधित होंगे, समाश्रयण निदर्श में X₂ को जोड़ने पर X₁ के समाश्रयण गुणांक का आकलक अपरिवर्तित रहेगा । 10 (d) प्रतिस्थापन रहित सरल यादृच्छिक प्रतिचयन द्वारा समष्टि की N इकाइयों से n आकार का एक प्रतिदर्श चुना गया । प्रतिस्थापन रहित सरल यादृच्छिक प्रतिचयन द्वारा n इकाइयों से n₁ इकाई का एक उप-प्रतिदर्श चुना गया । माना कि n₁ इकाइयों पर आधारित माध्य को ȳ₁ और n₂ = n - n₁ इकाइयों पर आधारित माध्य को ȳ₂ से व्यक्त किया गया । समष्टि माध्य Ȳₙ का आकलक दिया गया है : Ŷₙ = wȳ₁ + (1-w)ȳ₂ ; 0 < w < 1 दर्शाइए कि E(Ŷₙ) = Ȳₙ, और इसका प्रसरण प्राप्त कीजिए । 10 (e) किसी अभिकल्पना की दक्षता कैसे मापी जाती है ? पूर्णतः यादृच्छिकीकृत अभिकल्पना पर यादृच्छिकीकृत खंडक अभिकल्पना की दक्षता को मापने का व्यंजक व्युत्पन्न कीजिए। 10

Answer approach & key points

(a(i)) derive: given > assumptions > stepwise derivation > result > check | (a(ii)) derive: given > assumptions > stepwise derivation > result > check | (b) derive: given > assumptions > stepwise derivation > result > check | (c) derive: given > assumptions > stepwise derivation > result > check | (d) derive: given > assumptions > stepwise derivation > result > check | (e) derive: given > assumptions > stepwise derivation > result > check Full marks: Complete derivations with all steps, correct assumptions, and clear notation throughout.

  • State assumptions (e.g., E(ε)=0, Var(ε)=σ²)
  • Define Sum of Squared Errors (SSE) function
  • Differentiate SSE w.r.t β₀ and β₁
  • Solve normal equations for β̂₀ and β̂₁
  • Prove Σeᵢ = 0 using normal equations
  • Prove ΣXᵢeᵢ = 0 using normal equations
  • Show regression line passes through (X̄, Ȳ)
  • Derive ΣŶᵢeᵢ = 0 from previous results
Q6
50M derive Multivariate analysis, correlation, cluster sampling, multivariate normal distribution

(a) For a multiple linear regression model with three covariates X₁, X₂ and X₃, let rᵢⱼ denote the correlation coefficient between Xᵢ and Xⱼ. For a data, it was found r₁₂ = 0·77, r₂₃ = 0·52, r₁₃ = 0·72. (i) Check the consistency of the above data. (ii) If r₁₃ is unknown, obtain the limits within which r₁₃ lies given the above values for r₁₂ and r₂₃. 20 (b) In cluster sampling with equal size clusters, obtain the unbiased estimate of population mean. Also obtain its sampling variance as V(ȳ̄) = (1-f)(NM-1)S²{1+(M-1)ρcl}/[M²(N-1)n], where notations have their usual meanings. 15 (c) Let Z₃ₓ₁ = (X₁ₓ₁, Y₂ₓ₁)ᵀ ~ N₃((0, 0, 1)ᵀ, [[1, 2, 1], [2, 5, 2], [1, 2, 2]]). Show that conditional on X₁ₓ₁, the two components of Y₂ₓ₁ are independent but marginally they are not. 15

हिंदी में पढ़ें

(a) किसी बहु रैखिक समाश्रयण निदर्श जिसमें तीन सह-विचर X₁, X₂ और X₃ हैं, के लिए, माना rᵢⱼ, Xᵢ और Xⱼ में सहसंबंध गुणांक दर्शाता है। किन्हीं आँकड़ों के लिए, देखा गया कि r₁₂ = 0·77, r₂₃ = 0·52, r₁₃ = 0·72 है। (i) उपर्युक्त आँकड़ों की संगतता जाँचिए। (ii) यदि r₁₃ अज्ञात हो, तो ऊपर दिए गए r₁₂ और r₂₃ के मानों से r₁₃ की सीमाएँ प्राप्त कीजिए। 20 (b) समान आकार वाले गुच्छों के गुच्छ प्रतिचयन में, समष्टि माध्य का अनभिनत आकलक प्राप्त कीजिए। इसका प्रतिचयन प्रसरण भी निम्न रूप में ज्ञात कीजिए : V(ȳ̄) = (1-f)(NM-1)S²{1+(M-1)ρcl}/[M²(N-1)n] जहाँ संकेतों के अपने सामान्य अर्थ हैं। 15 (c) माना Z₃ₓ₁ = (X₁ₓ₁, Y₂ₓ₁)ᵀ ~ N₃((0, 0, 1)ᵀ, [[1, 2, 1], [2, 5, 2], [1, 2, 2]]). दर्शाइए कि X₁ₓ₁ के प्रतिबंध पर, Y₂ₓ₁ के दो घटक स्वतंत्र हैं लेकिन उपांतिय वे स्वतंत्र नहीं हैं। 15

Answer approach & key points

(a(i)) calculate: given > formula > substitution > result with units > interpretation | (a(ii)) calculate: given > formula > substitution > result with units > interpretation | (b) derive: given > assumptions > stepwise derivation > result > check | (c) calculate: given > formula > substitution > result with units > interpretation Full marks: All parts fully derived with correct notation and interpretation

  • Construct the 3x3 correlation matrix R
  • Calculate the determinant of R
  • Check if det(R) is non-negative
  • State the consistency condition for correlation coefficients
  • State the inequality for the determinant of a 3x3 correlation matrix
  • Substitute the known values r12 and r23
  • Solve the quadratic inequality for r13
  • State the final interval for r13
Q7
50M derive Factorial experiments, principal components, regression estimator

(a) (i) What is confounding in factorial experiments ? (ii) A 2^6factorial experiment is conducted in blocks of size2³. Write the confounded effects such that no main effect or two factor interaction are confounded. Give the list of independent and generalised interactions confounded along with the elements of key block only. (iii) Give the break-up of degrees of freedom for a 2^nfactorial experiment in2^k blocks. (b) What are principal components ? Describe how to compute the principal components of the vectors X₁ = 1 0 -1 and X₂ = -1 1 0 . Give X₁ and X₂ in terms of the principal components. (c) Define Regression estimator. Show bias = – Cov (x̄, b). Under what conditions is bias negligible ? Find the mean square error of the estimator to first degree of approximation. Give comparison of Regression estimator with Ratio estimator.

हिंदी में पढ़ें

(a) (i) बहु-उपादानी प्रयोगों में संकरण क्या है ? (ii) एक 2⁶बहु-उपादानी प्रयोग2³ आकार के खंडकों में संचालित किया गया। संकीर्ण प्रभावों को लिखिए जिसमें कोई भी मुख्य उपादान या दो घटक अन्योन्यक्रिया संकीर्ण न हों। संकीर्ण होने वाले स्वतंत्र व व्यापकीकृत अन्योन्यक्रियाओं की सूची लिखिए, साथ ही केवल प्रमुख खंडक के अवयव लिखिए। (iii) 2^kखंडकों में2ⁿ बहु-उपादानी प्रयोग के लिए स्वातंत्र्य कोटियों का विभाजन दीजिए। (b) मुख्य घटक क्या हैं ? सदिश X₁ = 1 0 -1 और X₂ = -1 1 0 के मुख्य घटकों के परिकलन का विवरण दीजिए । X₁ और X₂ को मुख्य घटकों के रूप में लिखिए । (c) समाश्रयण आकलक परिभाषित कीजिए । दर्शाइए अभिनति = – सहप्रसरण (x̄, b) । किन प्रतिबंधों के अंतर्गत अभिनति नगण्य होती है ? प्रथम घात के सन्निकट आकलक की त्रुटि वर्ग माध्य ज्ञात कीजिए । समाश्रयण आकलक की अनुपात आकलक के साथ तुलना कीजिए ।

Answer approach & key points

Framework: UPSC Statistics Paper 1. (a(i)) define: precise definition > the distinguishing feature > one example | (a(ii)) calculate: given > formula > substitution > result with units > interpretation | (a(iii)) enumerate: list the items in order > one line each > no commentary | (b) describe: define > structure or process in order > labelled diagram > significance | (c) derive: given > assumptions > stepwise derivation > result > check Full marks: All parts fully addressed with correct derivations, clear notation, and complete comparisons.

  • Define confounding as aliasing effects with blocks
  • Mention loss of independent estimation
  • State purpose: reduce experimental runs
  • Identify 3 independent interactions (3-factor)
  • List all 7 confounded effects (independent + generalized)
  • Calculate elements of the key block
  • Ensure no main or 2-factor effects are confounded
  • Total df = 2^n - 1
Q8
50M solve Stratified sampling, polynomial regression, split-plot designs

(a) (i) In stratified sampling under optimum allocation, how will you proceed to select units from different strata, if one or more nᵢ's happens to be greater than Nᵢ (i ≥ 2) ? (ii) A sample survey was conducted in a certain district of Himachal Pradesh. Four strata A, B, C and D of villages were formed according to the acreage of fruit trees as obtained from revenue records. A random sample of villages was selected from each stratum and the number of apple orchards in each selected village was noted. The data are shown below : | Stratum | Total number of villages (Nᵢ) | Number of villages in sample (nᵢ) | Number of orchards in the selected villages | |---------|------------------------------|-----------------------------------|---------------------------------------------| | A (0 – 3 acres) | 275 | 15 | 2, 5, 1, 9, 6, 7, 0, 4, 7, 0, 5, 0, 0, 3, 0 | | B (3 – 6 acres) | 146 | 10 | 21, 11, 7, 5, 6, 19, 5, 24, 30, 24 | | C (6 – 15 acres) | 93 | 12 | 3, 10, 4, 11, 38, 11, 4, 46, 4, 18, 1, 39 | | D (15 acres and above) | 62 | 11 | 30, 42, 20, 38, 29, 22, 31, 28, 66, 14, 15 | Estimate the number of orchards in the district. (b) (i) For a second order polynomial model with one predictor variable, derive the least squares normal equations clearly stating the conditions assumed. How will you interpret the parameters in this model ? (ii) Describe why it is recommended to work with predictor variables centred around the mean. Comment on fitted values of the response variable in this case. Prove your claim. (c) What are split-plot designs ? When do you recommend the use of such designs ? If e₁ and e₂ are the main plot and sub-plot errors respectively, both estimated in units of a single sub-plot, explain why e₁ is expected to be larger than e₂.

हिंदी में पढ़ें

(a) (i) स्तरीत प्रतिचयन में अनुकूलतम नियतन के अंतर्गत यदि एक या अधिक nᵢ, Nᵢ (i ≥ 2) से ज्यादा बड़े हैं, तो आप विभिन्न स्तरों से इकाइयों का चयन किस प्रकार करेंगे ? (ii) हिमाचल प्रदेश के किसी जिले में एक प्रतिदर्श सर्वेक्षण किया गया । राजस्व अभिलेखों द्वारा प्राप्त फलदार पेड़ों के क्षेत्रफल के आधार पर गाँवों के चार स्तर A, B, C और D बनाए गए । प्रत्येक स्तर से गाँवों का एक यादृच्छिक प्रतिदर्श चुना गया और प्रत्येक चुने गए गाँव से सेब के बगीचों की संख्या लिखी गई । आँकड़े नीचे दर्शाए गए हैं : | स्तर | गाँवों की कुल संख्या (Nᵢ) | प्रतिदर्श में गाँवों की संख्या (nᵢ) | चुने गए गाँवों में बगीचों की संख्या | |-----|------------------------|-------------------------------|--------------------------------| | A (0 – 3 एकड़) | 275 | 15 | 2, 5, 1, 9, 6, 7, 0, 4, 7, 0, 5, 0, 0, 3, 0 | | B (3 – 6 एकड़) | 146 | 10 | 21, 11, 7, 5, 6, 19, 5, 24, 30, 24 | | C (6 – 15 एकड़) | 93 | 12 | 3, 10, 4, 11, 38, 11, 4, 46, 4, 18, 1, 39 | | D (15 एकड़ और अधिक) | 62 | 11 | 30, 42, 20, 38, 29, 22, 31, 28, 66, 14, 15 | जिले में बगीचों की संख्या का आकलन कीजिए । (b) (i) द्विघातीय बहुपद निर्देश जिसमें एक प्रावकता चर है, के लिए माने गए प्रतिबंधों को स्पष्ट लिखते हुए, न्यूनतम वर्ग प्रसामान्य समीकरण व्युत्पन्न कीजिए । आप इस निर्देश में प्राचलों की व्याख्या कैसे करेंगे ? (ii) वर्णन कीजिए कि क्यों माध्य के परितः केंद्रित प्रावकता चरों को संस्तुत किया जाता है । इस विषय में अनुक्रिया चर के आसंगित मानों पर टिप्पणी लिखिए । अपने दावे को सिद्ध कीजिए । (c) विभक्त-क्षेत्र अभिकल्पनाएँ क्या हैं ? आप इन अभिकल्पनाओं के उपयोग को कब संस्तुत करेंगे ? यदि e₁ और e₂ क्रमशः मुख्य क्षेत्र और उप-क्षेत्र त्रुटियाँ हैं, दोनों ही एकल उप-क्षेत्र इकाइयों में आकलित हैं, तो स्पष्ट कीजिए कि क्यों e₁, e₂ से अधिक बड़ा अनुमानित होता है ।

Answer approach & key points

Framework: UPSC Statistics Paper 1. (a(i)) explain: definition/context > points in order > small example > short close | (a(ii)) calculate: given > formula > substitution > result with units > interpretation | (b(i)) derive: given > assumptions > stepwise derivation > result > check | (b(ii)) describe: define > structure or process in order > labelled diagram > significance | (c) explain: definition/context > points in order > small example > short close Full marks: Complete derivations with correct notation, accurate calculations, clear interpretations, and proper statistical reasoning throughout.

  • Identify strata where n_i > N_i
  • Specify use of census for those strata
  • Adjust allocation for remaining strata
  • Calculate sample mean for each stratum
  • Apply stratified estimator formula
  • Sum weighted stratum estimates
  • State final estimate clearly
  • State model y = β0 + β1x + β2x² + ε

Practice Statistics 2021 Paper I answer writing

Pick any question above, write your answer, and get a detailed AI evaluation against UPSC's standard rubric.

Start free evaluation →