Statistical Significance Tests for Machine Translation Evaluation

  title={Statistical Significance Tests for Machine Translation Evaluation},
  author={Philipp Koehn},
If two translation systems differ differ in performance on a test set, can we trust that this indicates a difference in true system quality? To answer this question, we describe bootstrap resampling methods to compute statistical significance of test results, and validate them on the concrete example of the BLEU score. Even for small test sizes of only 300 sentences, our methods may give us assurances that test result differences are real. 
Highly Influential
This paper has highly influenced 89 other papers. REVIEW HIGHLY INFLUENTIAL CITATIONS
Highly Cited
This paper has 1,208 citations. REVIEW CITATIONS
875 Citations
17 References
Similar Papers


Publications citing this paper.
Showing 1-10 of 875 extracted citations

1,208 Citations

Citations per Year
Semantic Scholar estimates that this publication has 1,208 citations based on the available data.

See our FAQ for additional information.

Similar Papers

Loading similar papers…