Q5 50M Compulsory derive Linear regression, multivariate normal, sample statistics, sample size, experimental design
(a) For a two variable linear regression model Yᵢ = a + bXᵢ + eᵢ, where E(eᵢ) = 0, Var(eᵢ) = σ²ₑ, Cov(eᵢ, eⱼ) = 0 for i ≠ j, (i,j) ∈ {1, 2, ..., n}, if â and b̂ are least square estimators of a and b respectively, derive expressions for Var(â), Var(b̂) and Cov(â, b̂). 10 marks
(b) Let X = (X₁ X₂ X₃)' ~ N₃(μ, Σ), where μ = (1 2 1)' and Σ = (9 2 2 / 2 3 0 / 2 0 2). Find the joint distribution of Y₁ = X₁ + X₂ + X₃ and Y₂ = X₂ - X₃. 10 marks
(c) If X₁, X₂, ..., Xₙ is a random sample from a standard normal population, then using quadratic forms show that the sample mean X̄ = (1/n)∑ⱼ₌₁ⁿ Xⱼ and sample variance S² = [1/(n-1)]∑ⱼ₌₁ⁿ(Xⱼ - X̄)² are stochastically independent. 10 marks
(d) Assume that in a population of very large number of items, proportion of defective items is 0·30. What should be the size of the sample, if a simple random sample is to be drawn from this population to estimate the percent defective within 2 percent of the true value with 95·5 percent probability? [Given P(0 ≤ Z ≤ 1·96) = 0·475; and P(0 ≤ Z ≤ 2·005) = 0·4775]. 10 marks
(e) How do the size and shape of plots and blocks effect the results of field experiments? 10 marks
हिंदी में पढ़ें
(a) द्विचर रैखिक समाश्रयन निदर्श Yᵢ = a + bXᵢ + eᵢ जहाँ E(eᵢ) = 0, Var(eᵢ) = σ²ₑ, Cov(eᵢ, eⱼ) = 0, i ≠ j, (i,j) ∈ {1, 2, ..., n}, के लिए यदि â और b̂ क्रमशः a और b के न्यूनतम वर्ग आकलक हैं, तो Var(â), Var(b̂) तथा Cov(â, b̂) के लिए व्यंजकों को व्युत्पन्न कीजिए। 10 अंक
(b) मान लीजिए X = (X₁ X₂ X₃)' ~ N₃(μ, Σ), जहाँ μ = (1 2 1)' तथा Σ = (9 2 2 / 2 3 0 / 2 0 2) है। Y₁ = X₁ + X₂ + X₃ और Y₂ = X₂ - X₃ का संयुक्त बंटन ज्ञात कीजिए। 10 अंक
(c) यदि X₁, X₂, ..., Xₙ एक मानक प्रसामान्य समष्टि से लिया गया एक यादृच्छिक प्रतिदर्श है, तो द्विघात रूपों का उपयोग करके दर्शाइए कि प्रतिदर्श माध्य X̄ = (1/n)∑ⱼ₌₁ⁿ Xⱼ और प्रतिदर्श प्रसरण S² = [1/(n-1)]∑ⱼ₌₁ⁿ(Xⱼ - X̄)² प्रसामान्य रूप से स्वतंत्र हैं। 10 अंक
(d) मान लें कि वस्तुओं की बहुत बड़ी संख्या वाली समष्टि में दोष पूर्ण वस्तुओं का अनुपात 0·30 है। इस समष्टि से एक सरल यादृच्छिक प्रतिदर्श निकाले जाने पर प्रतिदर्श का आमाप क्या होना चाहिए ताकि 95·5 प्रतिशत प्रायिकता के साथ वास्तविक मान के 2% के भीतर दोष प्रतिशत का आकलन किया जा सके? [दिया गया है P(0 ≤ Z ≤ 1·96) = 0·475; तथा P(0 ≤ Z ≤ 2·005) = 0·4775]। 10 अंक
(e) भूखंडों और खंडकों के आमाप और आकार खेत प्रयोगों के परिणामों को कैसे प्रभावित करते हैं? 10 अंक
Answer approach & key points
(a) derive: given > assumptions > stepwise derivation > result > check | (b) calculate: given > formula > substitution > result with units > interpretation | (c) derive: given > assumptions > stepwise derivation > result > check | (d) calculate: given > formula > substitution > result with units > interpretation | (e) discuss: intro > 3-4 dimensions > example > balanced close Full marks: Complete derivations with all steps, correct calculations, clear interpretation, proper notation throughout
- State OLS normal equations for â and b̂
- Express estimators as linear combinations of Y_i
- Apply Var and Cov properties to linear forms
- State final expressions for Var(â), Var(b̂), Cov(â, b̂)
- Define transformation matrix A for Y = AX
- Calculate mean vector μ_Y = Aμ
- Calculate covariance matrix Σ_Y = AΣA'
- State joint distribution as N₂(μ_Y, Σ_Y)
Q6 50M derive Bivariate normal, joint distributions, conditional expectation, generalized least squares
(a)(i) If (X, Y) follows bivariate normal BN(μ₁, μ₂, σ₁², σ₂², ρ), then obtain (A) E(e^X) (B) E(e^(X+Y)) (C) Var(e^X) and (D) Correlation between e^X and e^Y. 3+3+3+3=12 marks
(a)(ii) If (X, Y) have the joint probability density function g(x,y) = y e^(-y(x+1)), for x ≥ 0, y ≥ 0; 0 elsewhere, then find the regression curve of X on Y and comment on the nature of the curve. 8 marks
(b) Let X = (X₁, X₂, X₃)' ~ N₃(μ, Σ), in which μ = (2 1 3)' and Σ = (9 2 -2 / 2 2 -3 / -2 -3 9). Obtain (i) E{X₁ | X₂ = x₂, X₃ = x₃} and (ii) Var{X₁ | X₂ = x₂, X₃ = x₃}. 15 marks
(c) Consider the model: Y = X θ + ε, where ε is an n×1 vector of unobservable random variables such that E(ε) = 0 and D(ε) = σ²Ω, σ>0 unknown, Ω is a positive definite matrix of known constants and rank(X) = k<n. Then (i) Derive least square estimator of θ and (ii) Derive an unbiased estimator of σ². 9+6=15 marks
हिंदी में पढ़ें
(a)(i) यदि (X, Y) द्विचर प्रसामान्य BN(μ₁, μ₂, σ₁², σ₂², ρ) का अनुसरण करता है, तो (A) E(e^X) (B) E(e^(X+Y)) (C) Var(e^X) तथा (D) e^X और e^Y के बीच सहसंबंध ज्ञात कीजिए। 3+3+3+3=12 अंक
(a)(ii) यदि (X, Y) का संयुक्त प्रायिकता घनत्व फलन निम्नवत है: g(x,y) = y e^(-y(x+1)), x ≥ 0, y ≥ 0; 0 अन्यथा, तो X का Y पर समाश्रयन वक्र ज्ञात कीजिए तथा वक्र की प्रकृति पर टिप्पणी कीजिए। 8 अंक
(b) मान लीजिए कि X = (X₁, X₂, X₃)' ~ N₃(μ, Σ), जिसमें μ = (2 1 3)' तथा Σ = (9 2 -2 / 2 2 -3 / -2 -3 9) है। ज्ञात कीजिए (i) E{X₁ | X₂ = x₂, X₃ = x₃} और (ii) Var{X₁ | X₂ = x₂, X₃ = x₃}। 15 अंक
(c) निदर्श पर विचार कीजिए: Y = X θ + ε, जहाँ ε अलक्ष्य यादृच्छिक चरों का एक n×1 सदिश इस प्रकार है कि E(ε) = 0 और D(ε) = σ²Ω, σ > 0 अज्ञात है, Ω ज्ञात स्थिरांकों का एक धनात्मक निश्चित आव्यूह है तथा कोटि (X) = k < n है। तब: (i) θ का न्यूनतम वर्ग आकलक व्युत्पन्न कीजिए और (ii) σ² का एक अनभिनत आकलक व्युत्पन्न कीजिए। 9+6=15 अंक
Answer approach & key points
Framework: Generalized Least Squares (GLS) and Multivariate Normal Theory. (a) calculate: given > formula > substitution > result | (b) calculate: given > formula > substitution > result | (c) calculate: given > formula > substitution > result Full marks: Rigorous derivations with all matrix steps shown and correct final forms.
- MGF of bivariate normal
- Log-normal moments
- Marginal density integration
- Partitioned covariance matrix
- Conditional expectation formula
- Conditional variance formula
- GLS estimator derivation
- Unbiased variance estimator
Q7 50M analyse Latin square design and factorial experiments
(a) Analyse and interpret the following data concerning output of wheat per field obtained as a result of experiment conducted to test four varieties of wheat A, B, C and D under a Latin square design at 5% level of significance.
[Given F(3, 6) = 4·76; F(4, 7) = 4·12] (20 marks)
(b)(i) Explain the need of factorial experiments with an example from pharmaceutical study. (6 marks)
(b)(ii) Divide the 16 treatments of 2⁴ factorial experiment into 4 blocks of 4 treatments each, confounding the interaction effect AB and CD completely with blocks. Which other interaction is automatically confounded in this design ? (9 marks)
(c) Define Horvitz-Thompson estimator for estimating the population total, and show that it is unbiased for probability proportional to size sampling without replacement. Also find its sampling variance. (15 marks)
हिंदी में पढ़ें
(a) गेहूँ की चार किस्मों A, B, C और D के परीक्षण के लिए किये गये प्रयोग के परिणाम स्वरूप प्रति खेत गेहूँ के उत्पादन से संबंधित निम्नलिखित आँकड़ों का विश्लेषण और व्याख्या कीजिए, जो 5% सार्थकता स्तर पर एक लैटिन वर्ग अभिकल्पना के अंतर्गत किया गया हो ।
[दिया गया है F(3, 6) = 4·76; F(4, 7) = 4·12] (20 अंक)
(b)(i) बहु-उपादानी प्रयोगों की आवश्यकता की, एक औषध अध्ययन के उदाहरण के साथ, व्याख्या कीजिए। (6 अंक)
(b)(ii) 2⁴ बहु-उपादानी प्रयोग के 16 उपचारों को 4 समूहों में, प्रत्येक में 4 उपचारों के साथ, विभाजित कीजिए, जिसमें अन्योन्य क्रिया प्रभाव AB और CD को समूहों के साथ पूरी तरह से संकरण किया गया है। इस अभिकल्पना में कौन सी अन्य अन्योन्य क्रिया स्वचालित रूप से संकरित होती है ? (9 अंक)
(c) हारविट्ज-थॉम्पसन आकलक को समष्टि योग का आकलन करने के लिए परिभाषित कीजिए, और दर्शाइए कि यह आकार के समानुपात प्रायिकता वाले प्रतिचयन, प्रतिस्थापन रहित, के लिए अनभिनत है। इस का प्रतिचयन प्रसरण भी ज्ञात कीजिए। (15 अंक)
Answer approach & key points
Framework: UPSC Statistics Paper 1. (a) analyse: intro > causes > effects > stakeholders/linkages > way forward | (b(i)) explain: definition/context > points in order > small example > short close | (b(ii)) calculate: given > formula > substitution > result with units > interpretation | (c) derive: given > assumptions > stepwise derivation > result > check Full marks: Flawless ANOVA, correct block design, rigorous HT derivation
- Compute row, column, and treatment totals
- Calculate Sum of Squares for all sources
- Construct ANOVA table with F-ratios
- Compare F-values with given critical values
- Define factorial experiment concept
- Highlight advantage of studying interactions
- Provide specific pharmaceutical example
- Mention efficiency over one-factor-at-a-time
Q8 50M solve Principal components and missing value analysis
(a)(i) What are principal components ? Show that the principal components are uncorrelated. (10 marks)
(a)(ii) Obtain the principal components and the amount of variation explained by each principal component associated with the following dispersion matrix :
Σ = 4 & 2 & 1
2 & 3 & 1
1 & 1 & 2
Comment on the results. (10 marks)
(b) For the given data, the yield of the treatment B in the second block is missing and is denoted as 'y'. Estimate the missing value, and analyse the data by assuming the level of significance = 0·05.
[Given that F(3, 4) = 6·59; and F(2, 3) = 9·55] (20 marks)
(c) Distinguish between Sampling and Non-sampling Errors. What are their sources ? How these errors can be controlled ? (10 marks)
हिंदी में पढ़ें
(a)(i) मुख्य घटक क्या हैं ? दर्शाइए कि मुख्य घटक असहसंबंधित हैं। (10 अंक)
(a)(ii) निम्नलिखित प्रकीर्णन आव्यूह से संबंधित मुख्य घटकों को प्राप्त कीजिए तथा प्रत्येक मुख्य घटक द्वारा स्पष्ट की गई परिवर्तन की मात्रा प्राप्त कीजिए :
Σ = 4 & 2 & 1
2 & 3 & 1
1 & 1 & 2
परिणामों पर टिप्पणी कीजिए। (10 अंक)
(b) दिए गए आंकड़ों के लिए, दूसरे खंड में उपचार B की उपज लुप्त है और इसे 'y' से दर्शाया गया है। लुप्त मान का आकलन कीजिए, और आंकड़ों का सार्थकता स्तर 0·05 पर विश्लेषण कीजिए।
[दिया गया है F(3, 4) = 6·59; और F(2, 3) = 9·55] (20 अंक)
(c) प्रतिचयन और अप्रतिचयन त्रुटियों के बीच अंतर कीजिए। उनके स्रोत क्या हैं ? इन त्रुटियों को कैसे नियंत्रित किया जा सकता है ? (10 अंक)
Answer approach & key points
Framework: UPSC Statistics Paper 1. (a(i)) define: precise definition > the distinguishing feature > one example | (a(ii)) calculate: given > formula > substitution > result with units > interpretation | (b) calculate: given > formula > substitution > result with units > interpretation | (c) compare: paired headings or table > key differences > significance > conclusion Full marks: Rigorous derivation in (a), precise ANOVA in (b), clear distinction in (c).
- Define PC as linear combinations of variables
- State objective: maximize variance
- Show Cov(PC_i, PC_j) = 0 for i ≠ j
- Use eigenvalue properties of dispersion matrix
- Find eigenvalues of the 3x3 matrix
- Find corresponding eigenvectors
- Calculate proportion of variance for each PC
- Comment on the results (e.g., cumulative variance)