Skip to search formSkip to main contentSkip to account menu

Text corpus

Known as: Text corpora, Linguistic corpus, Text item 
In linguistics, a corpus (plural corpora) or text corpus is a large and structured set of texts (nowadays usually electronically stored and processed… 
Wikipedia (opens in a new tab)

Papers overview

Semantic Scholar uses AI to extract papers important to this topic.
Highly Cited
2015
Highly Cited
2015
Computational Paralinguistics has several unresolved issues, one of which is coping with large variability due to speakers… 
Review
2014
Review
2014
  • 2014
  • Corpus ID: 264700243
Foreword | The volume of digital forensic evidence is rapidly increasing, leading to large backlogs. In this paper, a Digital… 
Highly Cited
2010
Highly Cited
2010
We present a novel scheme to apply factored phrase-based SMT to a language pair with very disparate morphological structures. Our… 
Highly Cited
2005
Highly Cited
2005
Corpora for training plan recognizers are scarce and difficult to gather from humans. However, corpora could be a boon to plan… 
2004
2004
We deal with the automated acquisition of a Spanish medical subword lexicon from an already existing Portuguese seed lexicon… 
2004
2004
This paper presents a computational model of the incremental construction of an associative network from a corpus. It is aimed at… 
Highly Cited
2004
Highly Cited
2004
This paper presents the clustering algorithm PoBOC (Pole-Based Overlapping Clustering). It has two main characteristics: the… 
Highly Cited
2000
Highly Cited
2000
The Center for Integrated Acoustic Information Research (CIAIR) at Nagoya University has been collecting speech corpora in moving… 
Highly Cited
1998
Highly Cited
1998
This paper describes eight telephone-speech corpora at various stages of development at the Center for Spoken Language… 
Highly Cited
1960