Statistics

UPSC Statistics 2023 — Paper I

All 8 questions from UPSC Civil Services Mains Statistics 2023 Paper I (400 marks total). Every stem reproduced in full, with directive-word analysis, marks, word limits, and answer-approach pointers.

8Questions
400Total marks
2023Year
Paper IPaper

Topics covered

Probability theory and distributions (1)Joint distributions and convergence of random variables (1)Probability theory and statistical inference (1)Statistical inference and hypothesis testing (1)Multivariate normal distribution and linear models (1)Multivariate analysis and principal components (1)Sampling methods and stratified random sampling (1)Experimental design and statistical models (1)

A

Q1
50M Compulsory solve Probability theory and distributions

(a) Out of 1000 persons born, only 900 reach the age of 15 years, and out of every 1000 who reach the age of 15 years, 950 reach the age of 50 years. Out of every 1000 who reach the age of 50 years, 40 die in one year. Accordingly, what is the probability that a person would attain the age of 51 years ? (10 marks) (b) Let X be a continuous random variable with probability density function : f(x) = x/2, & 0 ≤ x < 1 1/2, & 1 ≤ x < 2 (3-x)/2, & 2 ≤ x < 3 0, & elsewhere Obtain the cumulative distribution function of X and hence find the value of P(X > 3/2). (10 marks) (c) Let Xₙ, n ≥ 1 be a sequence of mutually independent random variables such that P(Xₙ = nᵅ) = P(Xₙ = – nᵅ) = 0·5, for any α > 0. Derive the condition on α under which the sequence Xₙ, n ≥ 1 obeys WLLNs. (10 marks) (d) Apply Run Test to test the randomness of the following sequence of H and T at 5% level of significance : HHHHHHTHHHHHTHTHHHH TTHHHHTHHHTTHHHHHH THHTTHHTHHH Given : Z₍₀·₀₂₅₎ = 1·96 Z₍₀·₀₅₎ = 1·645 (10 marks) (e) Differentiate between prior and posterior distributions. In case of squared error loss function, find out the Bayes estimator for unknown parameter. (10 marks)

हिंदी में पढ़ें

(a) 1000 जन्म लेने वाले व्यक्तियों में से, केवल 900, 15 वर्ष तक की आयु तक पहुँच पाते हैं, तथा प्रति 1000 व्यक्ति जो 15 वर्ष की आयु तक पहुँचते हैं, उनमें से 950 व्यक्ति 50 वर्ष की आयु तक पहुँचते हैं। प्रति 1000 व्यक्तियों में जो 50 वर्ष की आयु तक पहुँचते हैं, उनमें से 40 व्यक्तियों की एक वर्ष में मृत्यु हो जाती है। तदनुसार एक व्यक्ति के 51 वर्ष की आयु तक पहुँचने की प्रायिकता क्या है ? (10 अंक) (b) माना X एक सतत यादृच्छिक चर है जिसका प्रायिकता घनत्व फलन है : f(x) = x/2, & 0 ≤ x < 1 1/2, & 1 ≤ x < 2 (3-x)/2, & 2 ≤ x < 3 0, & अन्यथा X का संचयी वितरण फलन निकालिए तथा इससे P(X > 3/2) का मान ज्ञात कीजिए। (10 अंक) (c) माना Xₙ, n ≥ 1 परस्पर स्वतंत्र यादृच्छिक चरों की श्रृंखला इस प्रकार है कि P(Xₙ = nᵅ) = P(Xₙ = – nᵅ) = 0·5, किसी भी α > 0 के लिए । α पर उस प्रतिबंध को निकालिए जिसके अंतर्गत, श्रृंखला Xₙ, n ≥ 1 निबल बृहत् संख्याओं के नियम (WLLNs) का पालन करती है । (10 अंक) (d) H एवं T के निम्नलिखित अनुक्रम की यादृच्छिकता जाँचने के लिए परम्परा (रन) परीक्षण, 5% सार्थकता स्तर, पर प्रयुक्त कीजिए : HHHHHHTHHHHHTHTHHHH TTHHHHTHHHTTHHHHHH THHTTHHTHHH दिया गया है : Z₍₀·₀₂₅₎ = 1·96 Z₍₀·₀₅₎ = 1·645 (10 अंक) (e) पूर्व एवं पश्च बंटनों में विभेद कीजिए । वर्ग-त्रुटि हानि फलन की स्थिति में अज्ञात प्राचल का बेज़ आकलक ज्ञात कीजिए । (10 अंक)

Answer approach & key points

Framework: UPSC Statistics Paper 1. (a) calculate: given > formula > substitution > result with units > interpretation | (b) derive: given > assumptions > stepwise derivation > result > check | (c) derive: given > assumptions > stepwise derivation > result > check | (d) calculate: given > formula > substitution > result with units > interpretation | (e) explain: definition/context > points in order > small example > short close Full marks: All parts answered with correct calculations, clear derivations, and proper interpretations.

  • Define survival probabilities for each age interval
  • Apply multiplication rule for sequential survival
  • Calculate probability of surviving from 50 to 51
  • Compute final product of all survival probabilities
  • Integrate the PDF to find the CDF for each interval
  • Ensure the CDF is continuous and non-decreasing
  • Calculate P(X > 3/2) using the CDF
  • Verify that the CDF approaches 1 as x approaches infinity
Q2
50M derive Joint distributions and convergence of random variables

(a) Let X, Y, Z be three mutually independent standard exponential variates and W₁ = X + Y + Z, W₂ = (X + Y)/(X + Y + Z), W₃ = X/(X + Y). Then (i) determine the joint distribution of W₁, W₂ and W₃. (ii) find out the marginal probability density functions of W₁, W₂ and W₃. (iii) examine the mutual independence of W₁, W₂ and W₃, and give your comment. (10+6+4=20 marks) (b) Give an example to prove or disprove the following : P(lim sup Aₙ) = 0 ⇒ Σₖ₌₁^∞ P(Aₖ) < ∞, for any sequence {Aₙ, n ≥ 1} of events defined on a probability space (Ω, 𝓐, P). (15 marks) (c) Let {Yₙ, n ≥ 1} be a sequence of random variables and Y be a degenerate random variable. Examine whether 'Yₙ converges in distribution to Y' implies 'Yₙ converges in probability to Y'. (15 marks)

हिंदी में पढ़ें

(a) माना X, Y, Z तीन परस्पर स्वतंत्र मानक घातीय चर हैं तथा W₁ = X + Y + Z, W₂ = (X + Y)/(X + Y + Z), W₃ = X/(X + Y). तब (i) W₁, W₂ एवं W₃ का संयुक्त बंटन निकालिए । (ii) W₁, W₂ एवं W₃ के सीमांत प्रायिकता घनत्व फलन ज्ञात कीजिए । (iii) W₁, W₂ एवं W₃ के परस्पर स्वतंत्र होने का परीक्षण कीजिए तथा इस पर अपनी टिप्पणी दीजिए । (10+6+4=20 अंक) (b) निम्नलिखित को सिद्ध या अस्वीकृत करने के लिए एक उदाहरण दीजिए : P(lim sup Aₙ) = 0 ⇒ Σₖ₌₁^∞ P(Aₖ) < ∞, जहाँ {Aₙ, n ≥ 1} घटनाओं की कोई श्रृंखला है जो कि संभाव्यता अंतराल (Ω, 𝓐, P) पर परिभाषित है । (15 अंक) (c) माना {Yₙ, n ≥ 1} यादृच्छिक चरों की एक श्रृंखला है तथा Y एक अपभ्रष्ट यादृच्छिक चर है । परीक्षण कीजिए कि क्या 'Yₙ बंटन में Y को अभिसरित होता है', से यह निष्कर्ष निकलता है कि 'Yₙ प्रायिकता में Y को अभिसरित होता है' । (15 अंक)

Answer approach & key points

(a(i)) derive: given > assumptions > stepwise derivation > result > check | (a(ii)) calculate: given > formula > substitution > result with units > interpretation | (a(iii)) examine: intro > how/why with reasoning > evidence > conclusion | (b) justify: claim > 3-4 reasons > evidence > conclusion | (c) examine: intro > how/why with reasoning > evidence > conclusion Full marks: Rigorous derivations, correct distributions, clear counterexamples, precise definitions.

  • Define inverse transformation X, Y, Z
  • Compute Jacobian determinant of transformation
  • Substitute into joint PDF of X, Y, Z
  • State support region for W1, W2, W3
  • Integrate joint PDF to find marginals
  • Identify distribution of W1 (Gamma)
  • Identify distribution of W2 (Beta)
  • Identify distribution of W3 (Beta)
Q3
50M prove Probability theory and statistical inference

(a) (i) If X is a random variable with finite variance, show that lim n² P{|X| > n} = 0. n → ∞ (10 marks) (ii) In a certain recruitment test, there are multiple choice questions. There are four possible options to each question, out of which one is correct. The probability of knowing correct option for an intelligent student is 90%, while it is 20% for a weaker student. An intelligent student ticks the correct option. What is the probability that he was guessing ? (10 marks) (b) Determine whether the sequence of mutually independent random variables {Xₙ, n ≥ 1}, in which P(Xₙ = ± n^λ) = 1/(2n^(2λ)) P(Xₙ = 0) = 1 - 1/n^(2λ) (λ < 1/2) obeys Central Limit Theorem (CLT) or not. (15 marks) (c) Define Sequential Probability Ratio Test (SPRT) along with its operating characteristic function and average sample number. Determine SPRT for testing H₀ : θ = 4 against H₁ : θ = 5 in N(θ, 1) with α = 0·5 and β = 0·2. (15 marks)

हिंदी में पढ़ें

(a) (i) यदि X परिमित प्रसरण के साथ एक यादृच्छिक चर है, तो दिखाइए कि lim n² P{|X| > n} = 0. n → ∞ (10 अंक) (ii) किसी एक भर्ती परीक्षा में, बहुविकल्पीय प्रश्न हैं । प्रत्येक प्रश्न में चार संभव विकल्प हैं, जिनमें से एक सही है । एक बुद्धिमान छात्र के सही विकल्प जानने की प्रायिकता 90% है, जबकि एक कमजोर छात्र की केवल 20% है । एक बुद्धिमान छात्र सही विकल्प पर निशान लगाता है । इसके अनुमान से सही विकल्प पर निशान लगाने की प्रायिकता क्या है ? (10 अंक) (b) परीक्षण कीजिए कि परस्पर स्वतंत्र यादृच्छिक चरों की श्रृंखला {Xₙ, n ≥ 1}, जिसमें P(Xₙ = ± n^λ) = 1/(2n^(2λ)) P(Xₙ = 0) = 1 - 1/n^(2λ) (λ < 1/2) केंद्रीय सीमा प्रमेय (CLT) का पालन करती है या नहीं । (15 अंक) (c) अनुक्रमिक प्रायिकता अनुपात परीक्षण (SPRT) को इसके संकारक अभिलक्षण फलन एवं औसत प्रतिदर्श संख्या के साथ परिभाषित कीजिए । N(θ, 1) में H₀ : θ = 4 विरुद्ध H₁ : θ = 5 के परीक्षण के लिए अनुक्रमिक प्रायिकता अनुपात परीक्षण (SPRT), α = 0·5 तथा β = 0·2 के साथ ज्ञात कीजिए । (15 अंक)

Answer approach & key points

(a(i)) derive: given > assumptions > stepwise derivation > result > check | (a(ii)) calculate: given > formula > substitution > result with units > interpretation | (b) examine: intro > how/why with reasoning > evidence > conclusion | (c) define: precise definition > the distinguishing feature > one example Full marks: Rigorous derivations, correct application of theorems, clear interpretation of results.

  • State assumption Var(X) < ∞
  • Apply Chebyshev's inequality or Markov's inequality
  • Show n²P{|X| > n} ≤ Var(X)
  • Conclude limit is 0
  • Define events: Intelligent (I), Weaker (W), Correct (C)
  • State P(C|I) = 0.9 and P(C|W) = 0.2
  • Apply Bayes' theorem formula
  • Calculate final probability value
Q4
50M explain Statistical inference and hypothesis testing

(a) What is the role of properties of completeness and sufficiency in Statistical Inference ? Explain. In U (0, θ), find out Uniformly Minimum Variance Unbiased Estimator (UMVUE) of θ. (20 marks) (b) A survey of 400 families with four children each have the following distribution : | Number of boys | 0 | 1 | 2 | 3 | 4 | |---|---|---|---|---|---| | Number of families | 16 | 89 | 145 | 118 | 32 | Is this result consistent with the hypothesis that male and female births are equally probable at 5% level of significance ? It is given that χ²_(.05) for 4 degrees of freedom = 9·488 and χ²_(.05) for 5 degrees of freedom = 11·070. (c) Define Likelihood Ratio Test. In N(θ, σ²), where σ² is unknown, find out LR test for testing H₀ : θ = θ₀ against H₁ : θ ∈ (Ω – θ₀), where Ω is the parametric space for θ. α is the size of the test.

हिंदी में पढ़ें

(a) सांख्यिकी निष्कर्ष में पूर्णता एवं पर्याप्तता के गुणों की क्या भूमिका है ? स्पष्ट कीजिए । U (0, θ) में, θ का एकसमान न्यूनतम प्रसरण अनभिनत आकलक (UMVUE) ज्ञात कीजिए । (20 अंक) (b) 400 परिवारों, जिनमें प्रत्येक में चार बच्चे हैं, के सर्वेक्षण का बंटन निम्नलिखित है : | लड़कों की संख्या | 0 | 1 | 2 | 3 | 4 | |---|---|---|---|---|---| | परिवारों की संख्या | 16 | 89 | 145 | 118 | 32 | क्या यह परिणाम 5% सार्थकता स्तर पर इस परिकल्पना से संगत है कि लड़कों एवं लड़कियों के जन्म होने की संभावना बराबर है ? यह दिया गया है कि χ²_(.०५) 4 स्वतंत्र कोटि के लिए = 9·488 एवं χ²_(.०५) 5 स्वतंत्र कोटि के लिए = 11·070. (c) संभाव्यता अनुपात परीक्षण को परिभाषित कीजिए । N(θ, σ²), जहाँ σ² अज्ञात है, में H₀ : θ = θ₀ विरुद्ध H₁ : θ ∈ (Ω – θ₀), जहाँ Ω, θ के लिए प्राचलिक अंतराल है, के परीक्षण के लिए संभाव्यता अनुपात (LR) परीक्षण ज्ञात कीजिए । α परीक्षण का आकार है ।

Answer approach & key points

(a) explain: definition/context > points in order > small example > short close | (b) calculate: given > formula > substitution > result with units > interpretation | (c) derive: given > assumptions > stepwise derivation > result > check Full marks: Rigorous derivations, correct test statistics, and clear interpretation of results.

  • Define completeness and sufficiency in inference
  • Identify a complete sufficient statistic for U(0, θ)
  • Apply Lehmann-Scheffé theorem for UMVUE
  • Derive the specific UMVUE of θ
  • State H0: p=0.5 and H1: p≠0.5
  • Calculate expected frequencies using Binomial distribution
  • Compute Chi-square statistic with correct degrees of freedom
  • Compare calculated value with critical value 9.488

B

Q5
50M Compulsory solve Multivariate normal distribution and linear models

(a) (i) If **X** = (X₁ X₂ X₃)' is distributed as N₃ (μ, Σ), find the distribution of [(X₁ – X₂) (X₂ – X₃)]'. (5 marks) (ii) Suppose that **X** = (X₁ X₂ X₃)' ~ N₃ (**0**, Σ), where Σ = 1 & ρ & 0 ρ & 1 & ρ 0 & ρ & 1 . Is there a value of ρ for which (X₁ + X₂ + X₃) and (X₁ – X₂ – X₃) are independent ? (5 marks) (b) Show that **X** = (X₁, X₂, ..., Xₚ)' has p-variate normal distribution if and only if every linear combination (l₁X₁ + l₂X₂ + ... + lₚXₚ) of **X** follows a univariate normal distribution. (10 marks) (c) Let x₁, x₂, ..., xₙ be n given observations, and suppose that Yᵢ = β₀ + β₁xᵢ + eᵢ; i = 1, 2, ..., n, where β₀, β₁ are unknown parameters and eᵢ are mutually independent normal random variables with E(eᵢ) = 0 and V(eᵢ) = σ², i = 1, 2, ..., n. Also, σ² is assumed to be unknown. Test the null hypothesis H₀ : β₀ = β₁ = 0. (10 marks) (d) Complete the following analysis of variance table of a design and examine whether there is a significant difference between the treatments at 5% level of significance: | Source of Variation | Degrees of Freedom | Sum of Squares | Mean Sum of Squares | Variance Ratio | |---------------------|-------------------|----------------|---------------------|----------------| | Blocks | — | 21 | 4·2 | — | | Treatments | — | — | 5·0 | — | | Error | 15 | 12 | — | | | Total | — | — | | | Given that F_·05(3, 15) = 8·70, F_·05(5, 15) = 4·62 (10 marks) (e) Define regression estimator used for the estimation of population mean. Obtain its bias and Mean Square Error (MSE) to the first order of approximation. (10 marks)

हिंदी में पढ़ें

(a) (i) यदि **X** = (X₁ X₂ X₃)' का बंटन N₃ (μ, Σ) है, तब [(X₁ – X₂) (X₂ – X₃)]' का बंटन ज्ञात कीजिए । (5 अंक) (ii) माना कि **X** = (X₁ X₂ X₃)' ~ N₃ (**0**, Σ) है, जहाँ Σ = 1 & ρ & 0 ρ & 1 & ρ 0 & ρ & 1 है । क्या ρ का ऐसा कोई मान है जिसके लिए (X₁ + X₂ + X₃) एवं (X₁ – X₂ – X₃) स्वतंत्र हैं ? (5 अंक) (b) दिखाइए कि **X** = (X₁, X₂, ..., Xₚ)' का बंटन p-चरिय प्रसामान्य बंटन है, यदि और केवल यदि **X** के प्रत्येक रैखीय युग्म (l₁X₁ + l₂X₂ + ... + lₚXₚ) का बंटन एकचरिय (एकविचर) प्रसामान्य बंटन है । (10 अंक) (c) माना x₁, x₂, ..., xₙ दिए हुए n प्रेक्षण हैं तथा Yᵢ = β₀ + β₁xᵢ + eᵢ; i = 1, 2, ..., n, जहाँ β₀, β₁ अज्ञात प्राचल हैं तथा सभी eᵢ E(eᵢ) = 0 एवं V(eᵢ) = σ², i = 1, 2, ..., n के साथ परस्पर स्वतंत्र प्रसामान्य यादृच्छिक चर हैं । σ² को अज्ञात माना गया है । निराकरणीय परिकल्पना H₀ : β₀ = β₁ = 0 का परीक्षण कीजिए । (10 अंक) (d) एक अभिकल्पना की निम्नलिखित प्रसरण विल्लेखन सारणी को पूर्ण कीजिए एवं 5% सार्थकता स्तर पर बताइए कि क्या व्यवहारों के मध्य सार्थक अंतर है : | विचरण स्रोत | स्वतंत्र कोटि | वर्गों का योग | माध्य वर्गों का योग | प्रसरण अनुपात | |------------|-------------|-------------|------------------|-------------| | खंड | — | 21 | 4·2 | — | | व्यवहार | — | — | 5·0 | — | | त्रुटि | 15 | 12 | — | | | योग | — | — | | | दिया गया है F_·05(3, 15) = 8·70, F_·05(5, 15) = 4·62 (10 अंक) (e) समष्टि माध्य के आकलन के लिए प्रयुक्त समाश्रयण आकलक को परिभाषित कीजिए । इसकी अभिनति (बायस) एवं माध्य वर्ग त्रुटि (एम.एस.ई.) को प्रथम सन्निकटन क्रम तक प्राप्त कीजिए । (10 अंक)

Answer approach & key points

Framework: Linear Model Theory and ANOVA. (a) calculate: given > formula > substitution > result with units > interpretation | (b) explain: definition/context > points in order > small example > short close | (c) calculate: given > formula > substitution > result with units > interpretation | (d) calculate: given > formula > substitution > result with units > interpretation | (e) explain: definition/context > points in order > small example > short close Full marks: Rigorous derivations, correct notation, complete tables, clear interpretations.

  • Multivariate normal transformation properties
  • Cramér-Wold theorem application
  • F-test for joint hypothesis in regression
  • ANOVA table completion and F-test
  • Regression estimator bias and MSE derivation
Q6
50M solve Multivariate analysis and principal components

(a) Let **X** = (X₁ X₂ X₃)' be distributed as N₃ (μ, Σ), where μ = (2 −1 3)' and Σ = 4 & 1 & 0 1 & 2 & 1 0 & 1 & 3 . Find (i) the conditional distribution of (X₁ X₂)' given X₃ = 2. (ii) partial correlation coefficient ρ₁₂.₃ and multiple correlation coefficient R₁.₂₃ (8+7 marks) (b) (i) Describe the complete analysis of two-way classified data with multiple (but equal) observations per cell, clearly stating the assumptions used. Also state two examples where such type of analysis is used. (ii) Let three mutually independent variables Y₁, Y₂ and Y₃ having common variance σ² and E(Y₁) = β₁ + β₂, E(Y₂) = β₁ + β₃, E(Y₃) = β₁ + β₂ be given. Show that the linear parametric function p₁β₁ + p₂β₂ + p₃β₃ is estimable if and only if p₁ = p₂ + p₃, clearly stating the assumptions used, if any. (5 marks) (c) (i) State briefly three reasons why an analyst may wish to perform a principal component analysis. (6 marks) (ii) Define canonical correlations and give two examples of their application. Describe the procedure of working out canonical correlations and canonical variates. (9 marks)

हिंदी में पढ़ें

(a) माना **X** = (X₁ X₂ X₃)' का बंटन N₃ (μ, Σ) है, जहाँ μ = (2 −1 3)' एवं Σ = 4 & 1 & 0 1 & 2 & 1 0 & 1 & 3 । ज्ञात कीजिए (i) (X₁ X₂)' का प्रतिबंधित बंटन जबकि X₃ = 2 दिया है । (ii) आंशिक सहसंबंध गुणांक ρ₁₂.₃ एवं बहु सहसंबंध गुणांक R₁.₂₃ (8+7 अंक) (b) (i) प्रति कोष्ठ संख्या में बराबर बहु आंकड़े (आब्जर्वेशन्स) रखने वाले द्वि-विध (टू-वे) वर्गीकृत आंकड़ों के सम्पूर्ण विश्लेषण का विवरण, उपयोग में ली गई मान्यताओं का स्पष्ट उल्लेख करते हुए दीजिए । ऐसे दो उदाहरण भी दीजिए जहाँ इस प्रकार के विश्लेषण का उपयोग होता है । (ii) माना कि तीन परस्पर स्वतंत्र चर Y₁, Y₂ और Y₃ जिनका प्रसरण σ² समान है तथा E(Y₁) = β₁ + β₂, E(Y₂) = β₁ + β₃, E(Y₃) = β₁ + β₂ दिए गए हैं । दिखाइए कि रैखीय प्राचलिक फलन p₁β₁ + p₂β₂ + p₃β₃ प्राकलिक है, यदि एवं केवल यदि p₁ = p₂ + p₃ है । साथ ही यदि कोई मान्यताएं प्रयुक्त होती हैं, तो उनका भी स्पष्ट उल्लेख कीजिए । (5 अंक) (c) (i) संक्षिप्त में तीन कारण लिखिए जिनके कारण विश्लेषक प्रमुख घटक विश्लेषण का प्रयोग करने की इच्छा कर सकता है। (6 अंक) (ii) विहित सहसंबंधों को परिभाषित कीजिए, तथा इनके अनुप्रयोग के दो उदाहरण दीजिए। विहित सहसंबंधों एवं विहित चरों को ज्ञात करने की विधि का वर्णन कीजिए। (9 अंक)

Answer approach & key points

Framework: UPSC Statistics Paper 1. (a(i)) calculate: given > formula > substitution > result with units > interpretation | (a(ii)) calculate: given > formula > substitution > result with units > interpretation | (b(i)) describe: define > structure or process in order > labelled diagram > significance | (b(ii)) derive: given > assumptions > stepwise derivation > result > check | (c(i)) highlight: name the salient points > one line of substance each > close | (c(ii)) define: precise definition > the distinguishing feature > one example Full marks: Rigorous derivations, correct matrix algebra, clear assumptions, and precise definitions.

  • Partition Σ into Σ11, Σ12, Σ21, Σ22
  • Compute conditional mean μ1.2 = μ1 + Σ12Σ22⁻¹(x2-μ2)
  • Compute conditional covariance Σ1.2 = Σ11 - Σ12Σ22⁻¹Σ21
  • State final N2 distribution with calculated values
  • Convert Σ to correlation matrix R
  • Apply formula for ρ12.3 using R elements
  • Apply formula for R1.23 using R elements
  • Provide final numerical values
Q7
50M discuss Sampling methods and stratified random sampling

(a) Discuss the difference between sampling for variables and sampling for attributes with examples. For a qualitative characteristic, find an unbiased estimator of population proportion along with its variance when sample is drawn by simple random sampling without replacement. Also obtain an unbiased estimator of this variance. 20 (b) The table given below gives the population and sample sizes, stratum means and variance of a stratified random sample of size 50. Symbols used have their usual meanings. | Stratum Number | Nᵢ | nᵢ | ȳᵢ | sᵢ² | |---|---|---|---|---| | 1 | 30 | 5 | 35 | 36 | | 2 | 50 | 10 | 40 | 49 | | 3 | 60 | 15 | 40 | 81 | | 4 | 60 | 20 | 55 | 144 | Verify that the existing allocation is optimum for given 4 strata. Also calculate the estimate of population variance under this allocation. 15 (c) Differentiate between Simple Random Sampling and Probability Proportional to Size Sampling. How will you draw a PPS sample of size n from a population of size N (n < N) by (i) Cumulative Total Method and (ii) Lahri's Method ? Explain. 15

हिंदी में पढ़ें

(a) चरों के प्रतिचयन एवं गुणात्मक चरों के प्रतिचयन में अंतर का उदाहरणों सहित वर्णन कीजिए। एक गुणात्मक अभिलक्षण के लिए समष्टि अनुपात का अनभिनत आकलक तथा इस आकलक का प्रसरण ज्ञात कीजिए जबकि प्रतिचयन प्रतिस्थापन रहित सरल यादृच्छिक विधि द्वारा किया गया है। इस प्रसरण का अनभिनत आकलक भी निकालिए। 20 (b) नीचे दी गई सारणी में 50 आकार के स्तरीकृत यादृच्छिक प्रतिदर्श के स्तरों का माध्य एवं प्रसरण तथा स्तरों की समष्टि का आकार तथा स्तरों से चयनित प्रतिदर्शी आकारों को दिया गया है। चिह्नों को उनके सामान्य अर्थों में प्रयुक्त किया गया है। | स्तर संख्या | Nᵢ | nᵢ | ȳᵢ | sᵢ² | |:---:|:---:|:---:|:---:|:---:| | 1 | 30 | 5 | 35 | 36 | | 2 | 50 | 10 | 40 | 49 | | 3 | 60 | 15 | 40 | 81 | | 4 | 60 | 20 | 55 | 144 | प्रमाणित कीजिए कि दिए गए 4 स्तरों के लिए मौजूदा आवंटन इष्टतम है। समष्टि प्रसरण का इस आवंटन के सापेक्ष आकलक भी ज्ञात कीजिए। 15 (c) सरल यादृच्छिक प्रतिचयन तथा आकार अनुपातिक प्रायिकता प्रतिचयन में विभेद कीजिए । एक n आकार के आकार अनुपातिक प्रायिकता प्रतिदर्श को आप (i) संचयी योग विधि तथा (ii) लाहिरी विधि द्वारा N (n < N) आकार की समष्टि से कैसे चुनेंगे ? स्पष्ट कीजिए । 15

Answer approach & key points

Framework: UPSC Statistics Paper 1. (a) discuss: intro > 3-4 dimensions > example > balanced close | (b) calculate: given > formula > substitution > result with units > interpretation | (c) explain: definition/context > points in order > small example > short close Full marks: All derivations complete, calculations accurate, methods clearly explained with correct notation

  • Define sampling for variables vs attributes with examples
  • Derive unbiased estimator of population proportion p
  • Derive variance of estimator under SRSWOR
  • Provide unbiased estimator of the variance
  • Verify existing allocation is optimum (Neyman allocation)
  • Calculate estimate of population variance
  • Show calculation steps clearly
  • Use correct stratified sampling formulas
Q8
50M differentiate Experimental design and statistical models

(a) Differentiate between randomised block design and balanced incomplete block design. In usual notations, for a balanced incomplete block design, prove that (i) bk = vr (ii) λ(v – 1) = r(k – 1) and (iii) b ≥ v. 20 (b) Explain the concept of confounding in design of experiment. In an experiment with three factors A, B and C, each at two levels, three replicates are divided in two blocks, each of four units. How will you confound ABC in the first, AC in the second and BC in the third replication ? 15 (c) Differentiate among fixed, random and mixed effect models with examples. How are the three basic principles of design fulfilled in randomised block design ? Explain. 15

हिंदी में पढ़ें

(a) यादृच्छिक खंड अभिकल्पना तथा संतुलित अपूर्ण खंडक अभिकल्पना में अंतर बताइए । सामान्य प्रयुक्त संकेताक्षरों में सिद्ध कीजिए कि संतुलित अपूर्ण खंडक अभिकल्पना में (i) bk = vr (ii) λ(v – 1) = r(k – 1) तथा (iii) b ≥ v. 20 (b) प्रयोगात्मक अभिकल्पना में संकरण के सिद्धांत की व्याख्या कीजिए । किसी प्रयोग में जिसमें तीन उपादान A, B तथा C जिनमें प्रत्येक दो स्तरों पर हैं, तीन पुनरावृत्त चार इकाइयों के दो खंडों में विभाजित हैं । आप ABC को पहले, AC को दूसरे तथा BC को तीसरे पुनरावृत्त में किस प्रकार संकीर्ण करेंगे ? 15 (c) नियत, यादृच्छिक एवं मिश्रित प्रभाव मॉडलों में उदाहरणों सहित विभेद कीजिए । यादृच्छिक खण्डक अभिकल्पना में अभिकल्पना के तीन मूलभूत सिद्धान्तों का समावेश कैसे होता है ? स्पष्ट कीजिए । 15

Answer approach & key points

Framework: Design of Experiments (DOE) & Statistical Models. (a) compare: paired headings or table > key differences > significance > conclusion | (b) explain: definition/context > points in order > small example > short close | (c) compare: paired headings or table > key differences > significance > conclusion Full marks: Rigorous proofs for (a); precise block construction for (b); clear model distinctions for (c).

  • Define RBD (complete blocks) vs BIBD (incomplete blocks)
  • Prove bk = vr via total treatment occurrences
  • Prove λ(v-1) = r(k-1) via pair co-occurrence
  • Prove b ≥ v using the derived relationship
  • Define confounding (aliasing with block effect)
  • Identify 4 treatment combinations for Block 1 (ABC)
  • Identify 4 treatment combinations for Block 2 (AC)
  • Identify 4 treatment combinations for Block 3 (BC)

Practice Statistics 2023 Paper I answer writing

Pick any question above, write your answer, and get a detailed AI evaluation against UPSC's standard rubric.

Start free evaluation →