Statistics 2021 Paper I 50 marks Construct

Paper I — Q3

(a) Let X and Y be two independent random variables following exponential distribution with mean 1/(λ) and 1/(μ) respectively, λ…

(a)

Let X and Y be two independent random variables following exponential distribution with mean 1/(λ) and 1/(μ) respectively, λ > 0, μ > 0. Suppose that (X₁, X₂, ..., Xₙ) and (Y₁, Y₂, ..., Yₙ) are sequences of observations on X and Y respectively. A random variable Uᵢ is defined as Uᵢ = 1, & if Xᵢ ≥ Yᵢ, i = 1, 2, ..., n 0, & otherwise Construct Wald's SPRT procedure based on Uᵢ's for testing H : λ = μ versus K : λ = 2μ with strength (α, β). 20 marks

(b)

Let Yᵢ, i ≥ 1 be independent and identical U(-1, 1) random variables. Determine if the following sequences converge in probability : (i) (Yᵢ)/i (ii) (Yᵢ)ⁱ (5+10 marks)

(c)

Let X₁, X₂, ..., Xₙ be a random sample from uniform distribution U(− θ, θ), θ > 0. Find the complete sufficient statistic for θ. Hence, obtain the best unbiased estimator of θ. 15 marks

हिंदी में प्रश्न पढ़ें
(a)

माना X और Y चर्यातांकी बंटन से लिए गए दो स्वतंत्र यादृच्छिक चर हैं जिनका माध्य क्रमशः 1/(λ) और 1/(μ), λ > 0, μ > 0 है । माना (X₁, X₂, ..., Xₙ) और (Y₁, Y₂, ..., Yₙ) क्रमशः X और Y से लिए गए प्रेक्षणों के अनुक्रम हैं । एक यादृच्छिक चर Uᵢ इस प्रकार से परिभाषित है Uᵢ = 1, & यदि Xᵢ ≥ Yᵢ, i = 1, 2, ..., n 0, & अन्यथा Uᵢ पर आधारित H : λ = μ विरुद्ध K : λ = 2μ के परीक्षण के लिए वाल्ड SPRT विधि की रचना कीजिए जिसकी शक्ति (α, β) है । (20 अंक)

(b)

माना Yᵢ, i ≥ 1, स्वतंत्र और सर्वसम U(-1, 1) यादृच्छिक चर हैं । ज्ञात कीजिए कि क्या निम्न अनुक्रम प्रायिकता में अभिसरित हैं : (i) (Yᵢ)/i (ii) (Yᵢ)ⁱ (5+10 अंक)

(c)

माना X₁, X₂, ..., Xₙ एकसमान बंटन U(− θ, θ), θ > 0 से लिया गया एक यादृच्छिक प्रतिदर्श है । θ का पूर्ण पर्याप्त प्रतिदर्शज्ञात कीजिए । इससे θ का सर्वोत्तम अनभिनत आकलक प्राप्त कीजिए । (15 अंक)

Q3 of the 2021 UPSC Mains Statistics Paper I, as printed
The question as printed in the 2021 Statistics paper

Model answer

Written by UPSC Answer Check against this question's marking rubric, to the expected length. UPSC does not publish answers for Mains — this is one way to score well, not an official key.

(a) Let p = P(X ≥ Y). Since X and Y are independent exponential with rates λ and μ,

p = ∫₀∞ P(X ≥ y) μ e^(−μy) dy = ∫₀∞ e^(−λy) μ e^(−μy) dy = μ/(λ+μ).

Ties have probability zero. Under H : λ = μ, p₀ = μ/(μ+μ) = 1/2. Under K : λ = 2μ, p₁ = μ/(2μ+μ) = 1/3. Hence U₁, U₂, ..., Uₙ are iid Bernoulli(p), with p₀ = 1/2 under H and p₁ = 1/3 under K. Let Sₙ = Σ Uᵢ.

For H : p = p₀ and K : p = p₁, the likelihood ratio is

Λₙ = ∏ [p₁^Uᵢ (1−p₁)^(1−Uᵢ)] / [p₀^Uᵢ (1−p₀)^(1−Uᵢ)] = (p₁/p₀)^Sₙ ((1−p₁)/(1−p₀))^(n−Sₙ) = (2/3)^Sₙ (4/3)^(n−Sₙ).

Therefore

log Λₙ = n log(4/3) − Sₙ log 2.

Let A = log(β/(1−α)) and B = log((1−β)/α), assuming 0 < α, β < 1 and α + β < 1. Wald’s SPRT is:

  • if log Λₙ ≤ A, stop and accept H;
  • if log Λₙ ≥ B, stop and reject H, i.e. accept K;
  • if A < log Λₙ < B, continue sampling.

Equivalently, in terms of Sₙ:

  • Accept H if Sₙ ≥ [n log(4/3) − log(β/(1−α))]/log 2.
  • Reject H if Sₙ ≤ [n log(4/3) − log((1−β)/α)]/log 2.
  • Continue if [n log(4/3) − log((1−β)/α)]/log 2 < Sₙ < [n log(4/3) − log(β/(1−α))]/log 2.

This is the required Wald SPRT based on the Uᵢ’s.

(b) Let Yᵢ ~ U(−1, 1) independently.

(i) For ε > 0,

P(|Yᵢ/i| > ε) = P(|Yᵢ| > iε).

Since |Yᵢ| ≤ 1 almost surely, this probability is 0 whenever iε ≥ 1. Hence for i > 1/ε, it is exactly 0. Therefore

P(|Yᵢ/i| > ε) → 0 as i → ∞.

Thus {Yᵢ/i} converges in probability to 0.

(ii) For ε > 0,

P(|Yᵢⁱ| > ε) = P(|Yᵢ| > ε^(1/i)).

If ε ≥ 1, this probability is 0. If 0 < ε < 1,

P(|Yᵢ| > ε^(1/i)) = 1 − ε^(1/i) = 1 − exp((log ε)/i).

As i → ∞, exp((log ε)/i) → 1, so the probability tends to 0. Hence {Yᵢⁱ} converges in probability to 0.

(c) The joint density of X₁, X₂, ..., Xₙ is

f(x₁, ..., xₙ; θ) = (2θ)^(−n), if −θ ≤ xᵢ ≤ θ for all i, and 0 otherwise.

This is equivalent to θ ≥ maxᵢ |xᵢ|. Therefore, by the factorization theorem, T = max₁≤i≤n |Xᵢ| is sufficient for θ.

The distribution of T is

P(T ≤ t) = P(|Xᵢ| ≤ t for all i) = (t/θ)ⁿ, 0 ≤ t ≤ θ.

Hence the density of T is

f_T(t) = n tⁿ⁻¹ / θⁿ, 0 < t < θ.

To prove completeness, suppose Eθ[g(T)] = 0 for all θ > 0. Then

∫₀θ g(t) n tⁿ⁻¹ / θⁿ dt = 0,

so

∫₀θ g(t) tⁿ⁻¹ dt = 0 for all θ > 0.

Differentiating with respect to θ gives g(θ)θⁿ⁻¹ = 0 almost everywhere, hence g = 0 almost everywhere. Thus T is complete sufficient.

Now

Eθ[T] = ∫₀θ t n tⁿ⁻¹ / θⁿ dt = nθ/(n+1).

Therefore

θ̂ = (n+1)/n T = (n+1)/n max₁≤i≤n |Xᵢ|

satisfies Eθ[θ̂] = θ. Since θ̂ is unbiased and is a function of the complete sufficient statistic T, by the Lehmann–Scheffé theorem it is the best unbiased estimator of θ.

What "Construct" is asking you to do

Build the required object — a velocity diagram, a sequential test, a control chart, a geometrical figure — step by step, so the sequence is visible on the page. The steps are marked, not only the finished thing.

Structure that answers it

Data and requirement → scale or basis chosen, stated → construction steps in order → the finished construction, labelled → quantities read off, or the result it yields

Where marks are lost

A diagram drawn without a stated scale, so nothing can be scaled off it and the quantities that follow lose their support. In statistics, writing down the procedure without fixing its defining constants — stopping bounds in terms of the two error probabilities, or the control limits — leaves it unmarkable.

All UPSC directive words, compared →

How this answer will be evaluated

Approach

Framework: Wald's Sequential Probability Ratio Test (SPRT). (a) calculate: given > formula > substitution > result with units > interpretation | (b) calculate: given > formula > substitution > result with units > interpretation | (c) calculate: given > formula > substitution > result with units > interpretation Full marks: Rigorous derivation with all steps shown and correct interpretation.

Key points expected

  • P(U_i=1) = lambda/(lambda+mu)
  • Likelihood ratio L_n = (lambda/(lambda+mu))^S * (mu/(lambda+mu))^(n-S)
  • Convergence in probability definition
  • M = max(|X_i|) is complete sufficient
  • UMVUE = (n+1)/n * M

Evaluation rubric

Each sub-part is marked on its own, against the marks and word limit printed on the paper.

  1. (a) Derive the likelihood ratio for U_i and define the stopping boundaries for H vs K. 20 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • Calculate P(U_i=1) under H and K
    • Formulate the likelihood ratio L_n
    • Define stopping boundaries A and B
    • State the decision rule for H, K, or continue

    Loses marks

    • Skipping the probability calculation for U_i
    • Confusing the likelihood ratio with the test statistic

    Earns more

    • Explicit calculation of P(X_i >= Y_i)
    • Correct substitution of lambda and mu values
    • Clear definition of alpha and beta roles

    Extra mark

    • Mention of expected sample size
  2. (b) Prove convergence in probability for the two given sequences. 15 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • Apply Markov's inequality or Chebyshev's inequality
    • Show limit of P(|X_n - c| > epsilon) is 0
    • Handle the bounded nature of U(-1, 1) for (i)
    • Analyze the exponential growth for (ii)

    Loses marks

    • Assuming convergence without proof
    • Incorrect application of limit laws

    Earns more

    • Explicit use of the definition of convergence in probability
    • Correct handling of the exponent in (ii)

    Extra mark

    • Alternative proof using almost sure convergence
  3. (c) Identify the complete sufficient statistic and derive the UMVUE for theta. 15 marks

    calculate— given → formula → substitution → result with units → interpretation

    Must cover

    • Identify M = max(|X_i|) as sufficient statistic
    • Prove completeness of M
    • Find an unbiased estimator based on M
    • Apply Lehmann-Scheffe theorem

    Loses marks

    • Using sum of X_i instead of max
    • Failing to prove completeness

    Earns more

    • Correct derivation of the distribution of M
    • Clear statement of the UMVUE formula

    Extra mark

    • Mention of the Cramer-Rao lower bound

Practice this exact question

Write your answer and it is marked point by point against the model answer above — what you covered, what you missed, what you got wrong.

Evaluate my answer →

More from Statistics 2021 Paper I