Statistics 2023 Paper I 50 marks Solve

Paper I — Q6

(a) Let **X** = (X₁ X₂ X₃)' be distributed as N₃ (μ, Σ), where μ = (2 −1 3)' and Σ = 4 & 1 & 0 1 & 2 & 1 0 & 1 & 3 . Find (i)…

(a)

Let **X** = (X₁ X₂ X₃)' be distributed as N₃ (μ, Σ), where μ = (2 −1 3)' and Σ = 4 & 1 & 0 1 & 2 & 1 0 & 1 & 3 . Find (i) the conditional distribution of (X₁ X₂)' given X₃ = 2. (ii) partial correlation coefficient ρ₁₂.₃ and multiple correlation coefficient R₁.₂₃ (8+7 marks)

(b)
(i)

Describe the complete analysis of two-way classified data with multiple (but equal) observations per cell, clearly stating the assumptions used. Also state two examples where such type of analysis is used. (ii) Let three mutually independent variables Y₁, Y₂ and Y₃ having common variance σ² and E(Y₁) = β₁ + β₂, E(Y₂) = β₁ + β₃, E(Y₃) = β₁ + β₂ be given. Show that the linear parametric function p₁β₁ + p₂β₂ + p₃β₃ is estimable if and only if p₁ = p₂ + p₃, clearly stating the assumptions used, if any. 5 marks

(c)
(i)

State briefly three reasons why an analyst may wish to perform a principal component analysis. (6 marks) (ii) Define canonical correlations and give two examples of their application. Describe the procedure of working out canonical correlations and canonical variates. 9 marks

हिंदी में प्रश्न पढ़ें
(a)

माना **X** = (X₁ X₂ X₃)' का बंटन N₃ (μ, Σ) है, जहाँ μ = (2 −1 3)' एवं Σ = 4 & 1 & 0 1 & 2 & 1 0 & 1 & 3 । ज्ञात कीजिए (i) (X₁ X₂)' का प्रतिबंधित बंटन जबकि X₃ = 2 दिया है । (ii) आंशिक सहसंबंध गुणांक ρ₁₂.₃ एवं बहु सहसंबंध गुणांक R₁.₂₃ (8+7 अंक)

(b)
(i)

प्रति कोष्ठ संख्या में बराबर बहु आंकड़े (आब्जर्वेशन्स) रखने वाले द्वि-विध (टू-वे) वर्गीकृत आंकड़ों के सम्पूर्ण विश्लेषण का विवरण, उपयोग में ली गई मान्यताओं का स्पष्ट उल्लेख करते हुए दीजिए । ऐसे दो उदाहरण भी दीजिए जहाँ इस प्रकार के विश्लेषण का उपयोग होता है । (ii) माना कि तीन परस्पर स्वतंत्र चर Y₁, Y₂ और Y₃ जिनका प्रसरण σ² समान है तथा E(Y₁) = β₁ + β₂, E(Y₂) = β₁ + β₃, E(Y₃) = β₁ + β₂ दिए गए हैं । दिखाइए कि रैखीय प्राचलिक फलन p₁β₁ + p₂β₂ + p₃β₃ प्राकलिक है, यदि एवं केवल यदि p₁ = p₂ + p₃ है । साथ ही यदि कोई मान्यताएं प्रयुक्त होती हैं, तो उनका भी स्पष्ट उल्लेख कीजिए । (5 अंक)

(c)
(i)

संक्षिप्त में तीन कारण लिखिए जिनके कारण विश्लेषक प्रमुख घटक विश्लेषण का प्रयोग करने की इच्छा कर सकता है। (6 अंक) (ii) विहित सहसंबंधों को परिभाषित कीजिए, तथा इनके अनुप्रयोग के दो उदाहरण दीजिए। विहित सहसंबंधों एवं विहित चरों को ज्ञात करने की विधि का वर्णन कीजिए। (9 अंक)

Q6 of the 2023 UPSC Mains Statistics Paper I, as printed
The question as printed in the 2023 Statistics paper

Model answer

Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.

(a) (i) Since X ~ N₃(μ,Σ), use the conditional normal formula. Partition X₁ = (X₁ X₂)' and X₂ = X₃. Here μ₁ = (2 −1)', μ₂ = 3, Σ₁₁ = ((4,1),(1,2)), Σ₁₂ = (0,1)', Σ₂₂ = 3. For X₃ = 2, E(X₁|X₃=2) = μ₁ + Σ₁₂Σ₂₂⁻¹(2−μ₂) = (2,−1)' + (0,1)'(1/3)(−1) = (2, −4/3)'. Cov(X₁|X₃=2) = Σ₁₁ − Σ₁₂Σ₂₂⁻¹Σ₂₁ = ((4,1),(1,2)) − (1/3)((0,0),(0,1)) = ((4,1),(1,5/3)). Hence (X₁,X₂)'|X₃=2 ~ N₂((2,−4/3)', ((4,1),(1,5/3))).

(a) (ii) The partial correlation is obtained from the conditional covariance above: ρ₁₂.₃ = 1/√(4 × 5/3) = √(3/20) = √15/10. For R₁.₂₃, use Σ₂₃ = ((2,1),(1,3)), cross-covariances (σ₁₂,σ₁₃) = (1,0). R₁.₂₃² = (1,0)Σ₂₃⁻¹(1,0)' / σ₁₁. Σ₂₃⁻¹ = (1/5)((3,−1),(−1,2)), so (1,0)Σ₂₃⁻¹(1,0)' = 3/5. Thus R₁.₂₃² = (3/5)/4 = 3/20, so ρ₁₂.₃ = √(3/20), R₁.₂₃ = √(3/20).

(b) (i) Let a factor A have a levels, factor B have b levels, and r equal replications per cell. Write y_ijk, i=1,…,a; j=1,…,b; k=1,…,r. The fixed-effects model is y_ijk = μ + α_i + β_j + γ_ij + ε_ijk, with Σα_i=0, Σβ_j=0, Σ_iγ_ij=0, Σ_jγ_ij=0, and ε_ijk ~ iid N(0,σ²). Assumptions: independent normal errors, constant variance σ², linear additive model with possible interaction, and fixed factor effects. Compute total SS = Σ(y_ijk−ȳ_...)², df = abr−1. Then SSA = brΣ_i(ȳ_i..−ȳ_...)², df = a−1; SSB = arΣ_j(ȳ_.j.−ȳ_...)², df = b−1; SSAB = rΣ_iΣ_j(ȳ_ij.−ȳ_i..−ȳ_.j.+ȳ_...)², df = (a−1)(b−1); SSE = Σ_iΣ_jΣ_k(y_ijk−ȳ_ij.)², df = ab(r−1). Test H₀A: all α_i=0 by F_A = MSA/MSE ~ F_a−1, ab(r−1); H₀B by F_B = MSB/MSE; H₀AB: all γ_ij=0 by F_AB = MSAB/MSE. If interaction is significant, interpret cell means/simple effects. If not, drop interaction and test main effects in the additive model. MSE estimates σ². Examples: crop yield under fertilizer levels and irrigation levels with replications; tensile strength under machines and operators with repeated measurements.

(b) (ii) Write Y = (Y₁,Y₂,Y₃)' and β=(β₁,β₂,β₃)'. From the mean model, E(Y) = Xβ, where X = ((1,1,0),(1,0,1),(1,1,0)). A linear function p'β is estimable iff p is in the column space of X', equivalently p = X'a for some a=(a₁,a₂,a₃)'. Now E(a'Y) = a₁(β₁+β₂)+a₂(β₁+β₃)+a₃(β₁+β₂) = (a₁+a₂+a₃)β₁ + (a₁+a₃)β₂ + a₂β₃. Thus p₁=a₁+a₂+a₃, p₂=a₁+a₃, p₃=a₂, giving p₁ = (a₁+a₃)+a₂ = p₂+p₃. Conversely, if p₁=p₂+p₃, choose a₂=p₃, a₁=p₂, a₃=0; then E(a'Y)=p'β. Hence p'β is estimable iff p₁ = p₂ + p₃. Assumptions: linear model, fixed β, independent errors with zero mean and common variance σ² for least-squares inference.

(c) (i) Three reasons for principal component analysis:

  • To reduce many correlated variables to a few uncorrelated principal components while retaining most variability.
  • To overcome multicollinearity or instability in regression by using orthogonal components.
  • To visualize/compress high-dimensional data and separate signal from noise.

(c) (ii) For two random vectors X (p×1) and Y (q×1), canonical correlations are the successive maximum correlations between linear combinations U=a'X and V=b'Y, with later pairs uncorrelated with earlier ones. Let Σ = ((Σxx,Σxy),(Σyx,Σyy)). The squared canonical correlations λ₁² ≥ … ≥ λ_k², k=min(p,q), are eigenvalues of Σxx⁻¹ΣxyΣyy⁻¹Σyx. The coefficient vector a satisfies (Σxx⁻¹ΣxyΣyy⁻¹Σyx − λ²I)a = 0, normalized by a'Σxxa=1. Then b ∝ Σyy⁻¹Σyxa, normalized by b'Σyyb=1; equivalently b = (1/λ)Σyy⁻¹Σyxa for λ>0. In practice replace Σ by sample covariance S, standardizing variables if scales differ. The canonical variates are U_i=a_i'X and V_i=b_i'Y. Applications: relating academic aptitude/achievement variables to job performance variables; relating economic indicators such as GDP and inflation to social indicators such as health and education.

What "Solve" is asking you to do

Choose the method, then carry it through to a final answer. Identifying what kind of problem this is and why that method applies is the first thing marked; a correct figure arrived at invisibly earns almost nothing.

Structure that answers it

Given data and what is required → method chosen, with the reason it applies → set-up (equation, circuit, free body, trial balance) → working, step by step → answer with units and any condition of validity

Where marks are lost

Doing the middle steps mentally and writing only the result. In mathematics papers, a further loss comes from giving a decimal where the exact value in surds or fractions was wanted, or from skipping the justification a part explicitly asks for.

All UPSC directive words, compared →

How this answer will be evaluated

Approach

Framework: UPSC Statistics Paper 1. (a(i)) calculate: given > formula > substitution > result with units > interpretation | (a(ii)) calculate: given > formula > substitution > result with units > interpretation | (b(i)) describe: define > structure or process in order > labelled diagram > significance | (b(ii)) derive: given > assumptions > stepwise derivation > result > check | (c(i)) highlight: name the salient points > one line of substance each > close | (c(ii)) define: precise definition > the distinguishing feature > one example Full marks: Rigorous derivations, correct matrix algebra, clear assumptions, and precise definitions.

Key points expected

  • Partition Σ into Σ11, Σ12, Σ21, Σ22
  • Compute conditional mean μ1.2 = μ1 + Σ12Σ22⁻¹(x2-μ2)
  • Compute conditional covariance Σ1.2 = Σ11 - Σ12Σ22⁻¹Σ21
  • State final N2 distribution with calculated values
  • Convert Σ to correlation matrix R
  • Apply formula for ρ12.3 using R elements
  • Apply formula for R1.23 using R elements
  • Provide final numerical values

Evaluation rubric

Each sub-part is marked on its own, against the marks and word limit printed on the paper.

  1. (a(i)) Conditional distribution of (X1, X2)' given X3=2. 8 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • Partition Σ into Σ11, Σ12, Σ21, Σ22
    • Compute conditional mean μ1.2 = μ1 + Σ12Σ22⁻¹(x2-μ2)
    • Compute conditional covariance Σ1.2 = Σ11 - Σ12Σ22⁻¹Σ21
    • State final N2 distribution with calculated values

    Loses marks

    • Incorrect partitioning of covariance matrix
    • Arithmetic errors in matrix multiplication

    Earns more

    • Correct matrix inversion for Σ22
    • Explicit substitution of x3=2

    Extra mark

    • Verification of positive definiteness
  2. (a(ii)) Partial correlation ρ12.3 and multiple correlation R1.23. 7 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • Convert Σ to correlation matrix R
    • Apply formula for ρ12.3 using R elements
    • Apply formula for R1.23 using R elements
    • Provide final numerical values

    Loses marks

    • Confusing covariance with correlation matrix
    • Sign errors in partial correlation formula

    Earns more

    • Correct identification of R11, R12, R13, R23

    Extra mark

    • Interpretation of correlation strength
  3. (b(i)) Analysis of two-way classified data with equal observations. 15 marks

    describe— define → structure or process in order → labelled diagram → significance

    Must cover

    • State model Yijk = μ + αi + βj + εijk
    • List assumptions (normality, independence, homoscedasticity)
    • Define ANOVA table components (SSA, SSB, SSAB, SSE)
    • State F-test hypotheses for main effects and interaction

    Loses marks

    • Omitting interaction term in model
    • Failing to state error term definition

    Earns more

    • Mention of balanced design benefits
    • Two specific examples (e.g., agriculture, quality control)

    Extra mark

    • Mention of post-hoc tests
  4. (b(ii)) Prove estimability condition p1 = p2 + p3. 5 marks

    derive— given → assumptions → stepwise derivation → result → check

    Must cover

    • Write expectation vector E(Y) = Xβ
    • Identify rank of design matrix X
    • Show p'β is estimable iff p is in row space of X
    • Derive p1 = p2 + p3 from linear dependence

    Loses marks

    • Assuming estimability without proof
    • Incorrect rank calculation

    Earns more

    • Explicit matrix representation of X

    Extra mark

    • General statement of estimability theorem
  5. (c(i)) Three reasons for performing PCA. 6 marks

    highlight— name the salient points → one line of substance each → close

    Must cover

    • Dimensionality reduction
    • Multicollinearity removal
    • Data visualization/structure identification

    Loses marks

    • Vague reasons like 'to analyze data'
    • Listing more than 3 without detail

    Earns more

    • Noise reduction
    • Feature extraction

    Extra mark

    • Mention of orthogonality of components
  6. (c(ii)) Definition, application, and procedure of canonical correlations. 9 marks

    define— precise definition → the distinguishing feature → one example

    Must cover

    • Define canonical correlations as max correlations of linear combos
    • Two application examples (e.g., psychology, economics)
    • Procedure: solve generalized eigenvalue problem
    • Define canonical variates as the linear combinations

    Loses marks

    • Confusing with principal components
    • Omitting the 'maximization' aspect in definition

    Earns more

    • Mention of SVD or matrix inversion in procedure
    • Interpretation of canonical loadings

    Extra mark

    • Mention of significance testing for correlations

Practice this exact question

Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.

Evaluate my answer →

More from Statistics 2023 Paper I