Statistics 2025 Paper II 50 marks Compulsory Explain

Paper II — Q5

(a) Explain the multicollinearity problem in a regression model. What are its consequences? State the different indicators of…

(a)

Explain the multicollinearity problem in a regression model. What are its consequences? State the different indicators of multicollinearity and explain. 10 marks

(b)

Establish the relationship among crude birthrate, general fertility rate and total fertility rate in the context of continuous data. Also, mention the properties of these fertility rates. 10 marks

(c)

What are the implications of using stable versus quasi-stable population assumption in demographic modelling? 10 marks

(d)

Discuss the problem of heteroscedasticity. Given that Yᵢ = α + β Xᵢ + Uᵢ with E(Uᵢ²) = K^2Xᵢ², prove that OLS estimates of α and β possess greater variance than OLS estimates of the transformed version of original model. 10 marks

(e)

What does it imply by validity of a test? Distinguish between the concepts of validity and reliability. 10 marks

हिंदी में प्रश्न पढ़ें
(a)

एक समाश्रयण निदर्श में बहुसंरेखता समस्या की व्याख्या कीजिए। इसके नतीजे क्या हैं? बहुसंरेखता के विभिन्न संकेतकों को बताइए तथा उनकी व्याख्या कीजिए। 10

(b)

संतत आँकड़ों के संदर्भ में अशोधित जनदर, सामान्य प्रजनन दर और कुल प्रजनन दर के बीच संबंध स्थापित कीजिए। इन प्रजनन दरों के गुणों का भी उल्लेख कीजिए। 10

(c)

जनसांख्यिकीय मॉडलिंग में स्थिर बनाम अर्ध-स्थिर जनसंख्या अनुमान का उपयोग करने के तात्पर्य क्या हैं? 10 marks

(d)

विषम विचलितता (हेटेरोस्केडेस्टिसिटी) की समस्या की विवेचना कीजिए। दिया गया है कि Yᵢ = α + β Xᵢ + Uᵢ साथ में E(Uᵢ²) = K^2Xᵢ², तो सिद्ध कीजिए कि α और β के साधारण न्यूनतम वर्ग (ओ० एल० एस०) आकलकों के प्रसरण, मूल मॉडल के रूपांतरित संस्करण के साधारण न्यूनतम वर्ग आकलकों के प्रसरण से अधिक हैं। 10

(e)

एक परीक्षण की वैधता से क्या अर्थ मिलता है? वैधता तथा विश्वसनीयता की अवधारणाओं के बीच का अंतर बताइए। 10

Q5 of the 2025 UPSC Mains Statistics Paper II, as printed
The question as printed in the 2025 Statistics paper

Model answer

Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.

The five issues link statistical inference with demographic measurement: diagnostics protect regression coefficients, while fertility and stability assumptions protect population projections.

(a) Multicollinearity. Multicollinearity occurs when regressors in a linear model are nearly linearly dependent, so one explanatory variable can be predicted with high R² from the others. It is distinct from simple pairwise correlation: with several regressors, near dependence can exist even when pairwise correlations are moderate. In applied Indian data, household income, education and family size in NSSO regressions may move together. It does not violate Gauss-Markov; OLS remains unbiased and consistent, but X′X is near singular. The causal chain is: near dependence creates small eigenvalues of X′X, inflates diagonal elements of (X′X)⁻¹, and raises Var(β̂ⱼ)=σ²[(X′X)⁻¹]ⱼⱼ. Large variances make coefficients unstable to small data changes and deflate t-statistics, so individually insignificant terms may appear even when the overall F-test is significant. Consequences include wide confidence intervals, sign reversals, and unreliable policy coefficients, because estimated partial effects may not be interpretable. Indicators are a high pairwise correlation matrix; auxiliary regressions of each Xⱼ on the others with high R²ⱼ; VIFⱼ=1/(1−R²ⱼ), where VIF>10 signals serious collinearity; a condition index, the square root of the ratio of largest to smallest eigenvalue of X′X, above 30 (severe above 1000); and high overall R² with insignificant t-ratios.

(b) Fertility rates. Let n(a) be the number of women of age a, P_F=∫₁₅⁴⁹ n(a)da, P total population, and B births in a year. If f(a) is the age-specific fertility rate, B=∫ f(a)n(a)da. Crude birth rate CBR=B/P; general fertility rate GFR=B/P_F. Hence CBR=GFR(P_F/P). Total fertility rate TFR=∫ f(a)da. Therefore GFR=[∫ f(a)n(a)da]/[∫ n(a)da]=TFR×R, where R=∫ f(a)n(a)da/[∫ n(a)da ∫ f(a)da] is an age-distribution adjustment factor. Thus CBR=TFR×R×(P_F/P). If women are uniformly distributed over the reproductive span L=34, R=1/L, so CBR≈(TFR/34)(P_F/P); the conversion factor is the span or weighted ratio, not the mean age of childbearing. The continuous formulation shows why CBR can fall even when TFR is unchanged if the share of reproductive-age women falls, as in ageing populations. Conversely, a young population can keep CBR high despite lower TFR. TFR is synthetic because it combines age-specific rates from one period, not the experience of one cohort. CBR and GFR are conventionally per 1000, TFR per woman. CBR is crude and age-structure dependent; GFR refines the denominator to reproductive-age women but still depends on the age mix within 15–49; TFR is age-standardized, a synthetic period measure, comparable across populations and useful for replacement-level analysis.

(c) Stable and quasi-stable populations. A stable population assumes constant age-specific fertility and mortality, and no net migration or a constant migration pattern. Lotka’s equation ∫ e⁻ʳᵃl(a)m(a)da=1 gives a constant growth rate r and a fixed age distribution proportional to e⁻ʳᵃl(a); the population grows exponentially. This simplifies projections, intercensal estimation, and separation of demographic momentum. But India’s fertility and mortality have changed over recent decades, so a strictly stable model can misstate growth and age structure. A quasi-stable population allows vital rates to change slowly; the age distribution is nearly stable and the growth rate changes gradually. This is closer to Census of India and UN medium-variant projections, where fertility decline, mortality improvement and migration are modelled over time. Quasi-stable assumptions are more realistic but more sensitive to trend assumptions and data quality. They capture momentum: even after fertility reaches replacement, a young age structure can sustain growth. Stable assumptions are useful as theoretical benchmarks; quasi-stable assumptions are preferred for planning, but both must be checked against Census and sample registration data.

(d) Heteroscedasticity. Heteroscedasticity means Var(Uᵢ) is not constant; here Var(Uᵢ)=K²Xᵢ². Assuming E(Uᵢ)=0, OLS remains unbiased but inefficient, and usual standard errors are invalid. Divide by Xᵢ (Xᵢ≠0): Yᵢ/Xᵢ=β+α(1/Xᵢ)+Uᵢ/Xᵢ. Let zᵢ=1/Xᵢ, yᵢ*=Yᵢ/Xᵢ, vᵢ=Uᵢ/Xᵢ; then Var(vᵢ)=K². OLS on this transformed model is WLS for the original model with weights 1/Xᵢ² and is BLUE. Its variance for β, the intercept, is K²/Σ(Xᵢ−X̄_w)²/Xᵢ², where X̄_w=Σ(Xᵢ/Xᵢ²)/Σ(1/Xᵢ²); its variance for α, the slope on zᵢ, is K²/Σ(zᵢ−z̄)², z̄=n⁻¹Σzᵢ. The transformed model downweights observations with large Xᵢ because their errors are more variable. The original OLS slope variance can be written as K²ΣXᵢ²(Xᵢ−X̄)²/Sxx², which is larger because it gives too much weight to high-variance observations. To prove the inequality, express original OLS in transformed variables. β_OLS=Σaᵢyᵢ*, aᵢ=Xᵢ(Xᵢ−X̄)/Sxx, Sxx=Σ(Xᵢ−X̄)². It is linear unbiased for β because Σaᵢ=1 and Σaᵢzᵢ=0. The transformed OLS intercept uses dᵢ=1/n−z̄(zᵢ−z̄)/Szz, Szz=Σ(zᵢ−z̄)², the minimum-norm vector satisfying the same constraints; since aᵢ−dᵢ is orthogonal to dᵢ, Σaᵢ²≥Σdᵢ². Hence Var(β_OLS)=K²Σaᵢ²≥K²Σdᵢ²=Var(β_WLS), strictly when Xᵢ vary. For α, α_OLS=Σcᵢyᵢ*, cᵢ=Xᵢ/n−X̄Xᵢ(Xᵢ−X̄)/Sxx. It satisfies Σcᵢ=0 and Σcᵢzᵢ=1. Cauchy-Schwarz gives 1=Σcᵢ(zᵢ−z̄)≤(Σcᵢ²)¹ᐟ²(Szz)¹ᐟ², so Var(α_OLS)=K²Σcᵢ²≥K²/Szz=Var(α_WLS). Thus OLS has greater variance.

(e) Validity and reliability. Validity is the degree to which a test measures what it claims to measure; it concerns systematic error and construct accuracy. Content validity asks whether items cover the domain, as in NSSO consumption surveys covering food, fuel, education and health. Criterion validity asks whether scores predict an external criterion, such as prelims scores predicting Mains performance. Construct validity asks whether the instrument captures the intended theoretical dimension, such as a poverty index capturing multidimensional deprivation. Reliability is precision or repeatability: consistency of scores under repeated measurement, assessed by test-retest, split-half or Cronbach’s alpha. Reliability estimates random error; validity estimates bias. A test can be reliable but invalid, consistently measuring the wrong thing; a valid test is generally reliable, but reliability alone does not establish validity. In survey design, reliability is improved by clear wording and trained enumerators, while validity requires correct sampling frames, measurement invariance and linkage to administrative records.

Together, these diagnostics ensure that regression and demographic models used in Indian planning produce robust, policy-relevant estimates.

What "Explain" is asking you to do

Make the working of something clear — what sets it off, what follows from what, and what it produces. Explain is the Commission's mechanism word: it dominates the technical papers and the “explain why” stems, where the marks sit in the causal chain and not in the label.

Structure that answers it

State what it is → the initiating condition → the chain of cause, step by step → an instance where it plays out → what the chain produces

Where marks are lost

Describing what something looks like instead of why it works that way. Naming the stages without linking them reads as description too.

All UPSC directive words, compared →

How this answer will be evaluated

Approach

Framework: UPSC Statistics Paper 2. (a) explain: definition/context > points in order > small example > short close | (b) explain: definition/context > points in order > small example > short close | (c) discuss: intro > 3-4 dimensions > example > balanced close | (d) derive: given > assumptions > stepwise derivation > result > check | (e) define: precise definition > the distinguishing feature > one example Full marks: Rigorous derivations, precise definitions, and clear distinctions.

Key points expected

  • Define multicollinearity as high correlation among regressors
  • List consequences: high variance, unstable coefficients
  • State indicators: VIF, correlation matrix, R²
  • Explain the mechanism of variance inflation
  • Define CBR, GFR, and TFR precisely
  • Establish the mathematical relationship among them
  • Mention properties: age-specific vs aggregate
  • Contextualize within continuous data

Evaluation rubric

Each sub-part is marked on its own, against the marks and word limit printed on the paper.

  1. (a) Definition, consequences, and indicators of multicollinearity in regression. 10 marks

    explain— definition/context → points in order → small example → short close

    Must cover

    • Define multicollinearity as high correlation among regressors
    • List consequences: high variance, unstable coefficients
    • State indicators: VIF, correlation matrix, R²
    • Explain the mechanism of variance inflation

    Loses marks

    • Confusing multicollinearity with perfect collinearity
    • Failing to distinguish from heteroscedasticity

    Earns more

    • Mention t-statistics become insignificant
    • Reference to condition number

    Extra mark

    • Provide a specific VIF threshold (e.g., >10)
  2. (b) Relationship and properties of CBR, GFR, and TFR. 10 marks

    explain— definition/context → points in order → small example → short close

    Must cover

    • Define CBR, GFR, and TFR precisely
    • Establish the mathematical relationship among them
    • Mention properties: age-specific vs aggregate
    • Contextualize within continuous data

    Loses marks

    • Using incorrect denominators in definitions
    • Failing to distinguish GFR from TFR

    Earns more

    • Note that TFR is a synthetic measure
    • Mention the denominator difference (total vs women)

    Extra mark

    • Provide a formula linking TFR to age-specific rates
  3. (c) Implications of stable vs quasi-stable population assumptions. 10 marks

    discuss— intro → 3-4 dimensions → example → balanced close

    Must cover

    • Define stable population assumption
    • Define quasi-stable population assumption
    • Compare implications for demographic modelling
    • Discuss the effect on projection accuracy

    Loses marks

    • Treating the two assumptions as identical
    • Ignoring the time dimension in stability

    Earns more

    • Mention the role of age structure
    • Reference to Lotka's equation

    Extra mark

    • Cite a specific demographic model (e.g., Leslie matrix)
  4. (d) Prove OLS variance is greater than transformed model variance. 10 marks

    derive— given → assumptions → stepwise derivation → result → check

    Must cover

    • State the heteroscedasticity condition E(Ui²)=K²Xi²
    • Derive the variance of OLS estimator for beta
    • Define the transformed model (divide by Xi)
    • Compare the variances to show OLS is larger

    Loses marks

    • Skipping the transformation step
    • Failing to compare the final variance terms

    Earns more

    • Show the algebraic steps clearly
    • Mention that transformed model is homoscedastic

    Extra mark

    • Explicitly write the variance of the transformed estimator
  5. (e) Define validity and distinguish it from reliability. 10 marks

    define— precise definition → the distinguishing feature → one example

    Must cover

    • Define validity of a test
    • Define reliability of a test
    • Distinguish between the two concepts
    • Explain the relationship (necessary vs sufficient)

    Loses marks

    • Confusing validity with accuracy
    • Failing to state that reliability is necessary for validity

    Earns more

    • Use the 'bullseye' analogy
    • Mention types of validity (content, criterion)

    Extra mark

    • Provide a specific example of a valid but unreliable test

Practice this exact question

Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.

Evaluate my answer →

More from Statistics 2025 Paper II