Paper I — Q5
(a) For a simple linear regression model Y = β₀ + β₁Xᵢ + εᵢ, i = 1, ..., n (i) Derive the least square estimators of β₀ and β₁…
For a simple linear regression model Y = β₀ + β₁Xᵢ + εᵢ, i = 1, ..., n
Derive the least square estimators of β₀ and β₁, clearly stating the conditions assumed.
For eᵢ = Yᵢ - Ŷᵢ where Ŷᵢ is the fitted value, show that 1. Σᵢ₌₁ⁿ eᵢ = 0 2. Σᵢ₌₁ⁿ Yᵢ = Σᵢ₌₁ⁿ Ŷᵢ 3. Σᵢ₌₁ⁿ Xᵢeᵢ = 0 4. Σᵢ₌₁ⁿ Ŷᵢeᵢ = 0 5. The regression line passes through (X̄, Ȳ). 5+5
In usual notations, if v, b, r, k and λ are the parameters of a Balanced Incomplete Block Design, then show that : b ≥ r + 1 ≥ λ + 2
v ≤ b ≤ (r² - 1)/λ 10
For the multiple linear regression model with two predictor variables X₁ and X₂, show that the estimate of regression coefficient of X₁ is unchanged when X₂ is added to the regression model, whenever X₁ and X₂ are uncorrelated. 10 marks
A sample of size n is drawn from a population having N units by simple random sampling without replacement. A sub-sample of n₁ units is drawn from the n units by simple random sampling without replacement. Let ȳ₁ denote the mean based on n₁ units and ȳ₂, the mean based on n₂ = n - n₁ units. Consider the estimator of the population mean Ȳₙ given by : Ŷₙ = wȳ₁ + (1-w)ȳ₂ ; 0 < w < 1 Show that E(Ŷₙ) = Ȳₙ, and obtain its variance. 10 marks
How is the efficiency of a design measured ? Derive the expression to measure the efficiency of a Randomised Block Design over a Completely Randomised Design. 10 marks
हिंदी में प्रश्न पढ़ें
एक साधारण रैखिक समाश्रयण निदर्श Y = β₀ + β₁Xᵢ + εᵢ, i = 1, ..., n के लिए
माने गए प्रतिबंधों को स्पष्ट लिखते हुए, β₀ और β₁ के न्यूनतम वर्ग आकलकों को व्युत्पन्न कीजिए।
eᵢ = Yᵢ - Ŷᵢ जहाँ Ŷᵢ आसंजित मान है, के लिए दर्शाइए कि 1. Σᵢ₌₁ⁿ eᵢ = 0 2. Σᵢ₌₁ⁿ Yᵢ = Σᵢ₌₁ⁿ Ŷᵢ 3. Σᵢ₌₁ⁿ Xᵢeᵢ = 0 4. Σᵢ₌₁ⁿ Ŷᵢeᵢ = 0 5. समाश्रयण रेखा (X̄, Ȳ) से गुजरती है। 5+5
प्रचलित संकेतों में, यदि v, b, r, k और λ किसी संतुलित अपूर्ण खंडक अभिकल्पना के प्राचल हैं, तो दर्शाइए कि : b ≥ r + 1 ≥ λ + 2
v ≤ b ≤ (r² - 1)/λ 10
एक बहु रैखिक समाश्रयण निदर्श जिसमें X₁ और X₂ दो प्रावकता चर हैं, के लिए दर्शाइए कि जब भी X₁ और X₂ असहसंबंधित होंगे, समाश्रयण निदर्श में X₂ को जोड़ने पर X₁ के समाश्रयण गुणांक का आकलक अपरिवर्तित रहेगा । 10
प्रतिस्थापन रहित सरल यादृच्छिक प्रतिचयन द्वारा समष्टि की N इकाइयों से n आकार का एक प्रतिदर्श चुना गया । प्रतिस्थापन रहित सरल यादृच्छिक प्रतिचयन द्वारा n इकाइयों से n₁ इकाई का एक उप-प्रतिदर्श चुना गया । माना कि n₁ इकाइयों पर आधारित माध्य को ȳ₁ और n₂ = n - n₁ इकाइयों पर आधारित माध्य को ȳ₂ से व्यक्त किया गया । समष्टि माध्य Ȳₙ का आकलक दिया गया है : Ŷₙ = wȳ₁ + (1-w)ȳ₂ ; 0 < w < 1 दर्शाइए कि E(Ŷₙ) = Ȳₙ, और इसका प्रसरण प्राप्त कीजिए । 10
किसी अभिकल्पना की दक्षता कैसे मापी जाती है ? पूर्णतः यादृच्छिकीकृत अभिकल्पना पर यादृच्छिकीकृत खंडक अभिकल्पना की दक्षता को मापने का व्यंजक व्युत्पन्न कीजिए। 10
The figure this question refers to, in words
The question paper is a scan and the diagram did not survive as text. This is the figure as read from the original page — every component, value and label — so the question can be worked from the text below.
Table: Chi-Square (chi-square) Distribution - Area to the Right of Critical Value. The table lists Degrees of Freedom (1 to 20) in the first column. The subsequent columns represent the area to the right of the critical value with headers: 0.995, 0.99, 0.975, 0.95, 0.90, 0.10, 0.05, 0.025, 0.01, 0.005. The values are as follows: Row 1: -, -, 0.001, 0.004, 0.016, 2.706, 3.841, 5.024, 6.635, 7.879. Row 2: 0.010, 0.020, 0.051, 0.103, 0.211, 4.605, 5.991, 7.378, 9.210, 10.597. Row 3: 0.072, 0.115, 0.216, 0.352, 0.584, 6.251, 7.815, 9.348, 11.345, 12.838. Row 4: 0.207, 0.297, 0.484, 0.711, 1.064, 7.779, 9.488, 11.143, 13.277, 14.860. Row 5: 0.412, 0.554, 0.831, 1.145, 1.610, 9.236, 11.071, 12.833, 15.086, 16.750. Row 6: 0.676, 0.872, 1.237, 1.635, 2.204, 10.645, 12.592, 14.449, 16.812, 18.548. Row 7: 0.989, 1.239, 1.690, 2.167, 2.833, 12.017, 14.067, 16.013, 18.475, 20.278. Row 8: 1.344, 1.646, 2.180, 2.733, 3.490, 13.362, 15.507, 17.535, 20.090, 21.955. Row 9: 1.735, 2.088, 2.700, 3.325, 4.168, 14.684, 16.919, 19.023, 21.666, 23.589. Row 10: 2.156, 2.558, 3.247, 3.940, 4.865, 15.987, 18.307, 20.483, 23.209, 25.188. Row 11: 2.603, 3.053, 3.816, 4.575, 5.578, 17.275, 19.675, 21.920, 24.725, 26.757. Row 12: 3.074, 3.571, 4.404, 5.226, 6.304, 18.549, 21.026, 23.337, 26.217, 28.299. Row 13: 3.565, 4.107, 5.009, 5.892, 7.042, 19.812, 22.362, 24.736, 27.688, 29.819. Row 14: 4.075, 4.660, 5.629, 6.571, 7.790, 21.064, 23.685, 26.119, 29.141, 31.319. Row 15: 4.601, 5.229, 6.262, 7.261, 8.547, 22.307, 24.996, 27.488, 30.578, 32.801. Row 16: 5.142, 5.812, 6.908, 7.962, 9.312, 23.542, 26.296, 28.845, 32.000, 34.267. Row 17: 5.697, 6.408, 7.564, 8.672, 10.085, 24.769, 27.587, 30.191, 33.409, 35.718. Row 18: 6.265, 7.015, 8.231, 9.390, 10.865, 25.989, 28.869, 31.526, 34.805, 37.156. Row 19: 6.844, 7.633, 8.907, 10.117, 11.651, 27.204, 30.144, 32.852, 36.191, 38.582. Row 20: 7.434, 8.260, 9.591, 10.851, 12.443, 28.412, 31.410, 34.170, 37.566, 39.997.
What "Derive" is asking you to do
Reach the stated expression from a starting relation, justifying every step. The destination is printed in the question, so only the route earns marks, and the assumptions you work under are part of that route.
Structure that answers it
Assumptions and notation defined → starting relation or governing equation → each step with its justification → the required expression → limiting case or boundary check
Where marks are lost
Writing the standard result first and fitting three lines to it, which an examiner reads at a glance. Marks also go on assumptions left unstated — lossless medium, small amplitude, errors independent with zero mean — and on symbols used before they are defined, even when the question says usual notations.
How this answer will be evaluated
Approach
(a(i)) derive: given > assumptions > stepwise derivation > result > check | (a(ii)) derive: given > assumptions > stepwise derivation > result > check | (b) derive: given > assumptions > stepwise derivation > result > check | (c) derive: given > assumptions > stepwise derivation > result > check | (d) derive: given > assumptions > stepwise derivation > result > check | (e) derive: given > assumptions > stepwise derivation > result > check Full marks: Complete derivations with all steps, correct assumptions, and clear notation throughout.
Key points expected
- State assumptions (e.g., E(ε)=0, Var(ε)=σ²)
- Define Sum of Squared Errors (SSE) function
- Differentiate SSE w.r.t β₀ and β₁
- Solve normal equations for β̂₀ and β̂₁
- Prove Σeᵢ = 0 using normal equations
- Prove ΣXᵢeᵢ = 0 using normal equations
- Show regression line passes through (X̄, Ȳ)
- Derive ΣŶᵢeᵢ = 0 from previous results
Evaluation rubric
Each sub-part is marked on its own, against the marks and word limit printed on the paper.
- (a(i)) Derive least square estimators for β₀ and β₁ with stated assumptions.
derive— given → assumptions → stepwise derivation → result → check
Must cover
- State assumptions (e.g., E(ε)=0, Var(ε)=σ²)
- Define Sum of Squared Errors (SSE) function
- Differentiate SSE w.r.t β₀ and β₁
- Solve normal equations for β̂₀ and β̂₁
Loses marks
- Derivation without stating assumptions
- Skipping the differentiation step
Earns more
- Explicitly write the normal equations
- Show β̂₁ = Sxy/Sxx and β̂₀ = Ȳ - β̂₁X̄
Extra mark
- Mention Gauss-Markov theorem context
- (a(ii)) Prove the five algebraic properties of residuals and fitted values.
derive— given → assumptions → stepwise derivation → result → check
Must cover
- Prove Σeᵢ = 0 using normal equations
- Prove ΣXᵢeᵢ = 0 using normal equations
- Show regression line passes through (X̄, Ȳ)
- Derive ΣŶᵢeᵢ = 0 from previous results
Loses marks
- Stating results without proof
- Confusing residuals with errors
Earns more
- Show ΣYᵢ = ΣŶᵢ as a direct consequence of Σeᵢ=0
- Clear step-by-step algebraic manipulation
Extra mark
- Geometric interpretation of orthogonality
- (b) Prove the two inequalities for Balanced Incomplete Block Design parameters. 10 marks
derive— given → assumptions → stepwise derivation → result → check
Must cover
- Use BIBD relations (vr=bk, λ(v-1)=r(k-1))
- Prove b ≥ r + 1 ≥ λ + 2
- Prove v ≤ b ≤ (r² - 1)/λ
- Clearly define v, b, r, k, λ
Loses marks
- Using incorrect BIBD relations
- Skipping algebraic steps in inequality proof
Earns more
- Logical flow from basic BIBD identities
- Correct algebraic manipulation of inequalities
Extra mark
- Mention Fisher's inequality context
- (c) Show regression coefficient of X₁ is unchanged when X₂ is added if uncorrelated. 10 marks
derive— given → assumptions → stepwise derivation → result → check
Must cover
- Write normal equations for simple regression (X₁ only)
- Write normal equations for multiple regression (X₁, X₂)
- Use condition Σ(X₁-X̄₁)(X₂-X̄₂) = 0
- Show β̂₁ is identical in both cases
Loses marks
- Not explicitly using the uncorrelated condition
- Confusing correlation with independence
Earns more
- Clear comparison of the two coefficient formulas
- Explicit use of uncorrelated condition
Extra mark
- Mention orthogonality of design matrix
- (d) Show estimator is unbiased and derive its variance. 10 marks
derive— given → assumptions → stepwise derivation → result → check
Must cover
- Show E(Ŷₙ) = Ȳₙ using linearity of expectation
- Derive Var(Ŷₙ) using variance of sample means
- Account for covariance between ȳ₁ and ȳ₂
- Use finite population correction factors
Loses marks
- Ignoring covariance between sub-samples
- Using wrong variance formula for SRSWOR
Earns more
- Correct expression for Var(ȳ₁) and Var(ȳ₂)
- Correct covariance term derivation
Extra mark
- Mention optimal w for minimum variance
- (e) Derive efficiency of Randomised Block Design over Completely Randomised Design. 10 marks
derive— given → assumptions → stepwise derivation → result → check
Must cover
- Define efficiency as ratio of variances
- Write Var(Ȳ) for CRD
- Write Var(Ȳ) for RBD
- Derive efficiency formula E = Var(CRD)/Var(RBD)
Loses marks
- Using incorrect variance formulas
- Not defining efficiency clearly
Earns more
- Correct variance expressions for both designs
- Clear definition of efficiency measure
Extra mark
- Mention conditions for RBD superiority
Model answer coming soon
Every evaluation on this site is marked against a verified model answer. This question's answer is still being written; evaluation opens the moment it lands.
More from Statistics 2021 Paper I
- Q2 (a) Let Y₁, Y₂, Y₃, ... be independent and identical Poisson random variables with parame…
- Q3 (a) Let X and Y be two independent random variables following exponential distribution wi…
- Q4 (a) Let X₁, X₂, ..., Xₙ be a random sample from Poisson distribution with mean λ > 0. Def…
- Q5 (a) For a simple linear regression model Y = β₀ + β₁Xᵢ + εᵢ, i = 1, ..., n (i) Derive the…
- Q6 (a) For a multiple linear regression model with three covariates X₁, X₂ and X₃, let rᵢⱼ d…
- Q7 (a) (i) What is confounding in factorial experiments ? (ii) A 2^6factorial experiment is…
- Q8 (a) (i) In stratified sampling under optimum allocation, how will you proceed to select u…