Statistical Significance Tests for Machine Translation Evaluation


If two translation systems differ differ in performance on a test set, can we trust that this indicates a difference in true system quality? To answer this question, we describe bootstrap resampling methods to compute statistical significance of test results, and validate them on the concrete example of the BLEU score. Even for small test sizes of only 300 sentences, our methods may give us assurances that test result differences are real.

Extracted Key Phrases

12 Figures and Tables

Citations per Year

1,062 Citations

Semantic Scholar estimates that this publication has 1,062 citations based on the available data.

See our FAQ for additional information.

Cite this paper

@inproceedings{Koehn2004StatisticalST, title={Statistical Significance Tests for Machine Translation Evaluation}, author={Philipp Koehn}, booktitle={EMNLP}, year={2004} }