Skip to search form
Skip to main content
Skip to account menu
Semantic Scholar
Semantic Scholar's Logo
Search 236,830,568 papers from all fields of science
Search
Sign In
Create Free Account
Text corpus
Known as:
Text corpora
, Linguistic corpus
, Text item
Expand
In linguistics, a corpus (plural corpora) or text corpus is a large and structured set of texts (nowadays usually electronically stored and processed…
Expand
Wikipedia
(opens in a new tab)
Create Alert
Alert
Related topics
Related topics
50 relations
Amarna letter EA 256
Amarna letter EA 365
Amarna letters–localities and their rulers
Amebis
Expand
Papers overview
Semantic Scholar uses AI to extract papers important to this topic.
Highly Cited
2015
Highly Cited
2015
Fisher vectors with cascaded normalization for paralinguistic analysis
Heysem Kaya
,
Alexey Karpov
,
A. A. Salah
Interspeech
2015
Corpus ID: 14524510
Computational Paralinguistics has several unresolved issues, one of which is coping with large variability due to speakers…
Expand
Review
2014
Review
2014
framework for digital forensic evidence : Storage , intelligence , review and archive
2014
Corpus ID: 264700243
Foreword | The volume of digital forensic evidence is rapidly increasing, leading to large backlogs. In this paper, a Digital…
Expand
Highly Cited
2010
Highly Cited
2010
Syntax-to-Morphology Mapping in Factored Phrase-Based Statistical Machine Translation from English to Turkish
Reyyan Yeniterzi
,
Kemal Oflazer
Annual Meeting of the Association for…
2010
Corpus ID: 14292100
We present a novel scheme to apply factored phrase-based SMT to a language pair with very disparate morphological structures. Our…
Expand
Highly Cited
2005
Highly Cited
2005
Generating Artificial Corpora for Plan Recognition
Nate Blaylock
,
James F. Allen
User Modeling
2005
Corpus ID: 10500675
Corpora for training plan recognizers are scarce and difficult to gather from humans. However, corpora could be a boon to plan…
Expand
2004
2004
Cognate Mapping - A Heuristic Strategy for the Semi-Supervised Acquisition of a Spanish Lexicon from a Portuguese Seed Lexicon
S. Schulz
,
K. Markó
,
E. Sbrissia
,
P. Nohama
,
U. Hahn
International Conference on Computational…
2004
Corpus ID: 4439397
We deal with the automated acquisition of a Spanish medical subword lexicon from an already existing Portuguese seed lexicon…
Expand
2004
2004
Incremental Construction of an Associative Network from a Corpus
Benoît Lemaire
,
G. Denhière
2004
Corpus ID: 8933441
This paper presents a computational model of the incremental construction of an associative network from a corpus. It is aimed at…
Expand
Highly Cited
2004
Highly Cited
2004
PoBOC: An Overlapping Clustering Algorithm, Application to Rule-Based Classification and Textual Data
Guillaume Cleuziou
,
Lionel Martin
,
Christel Vrain
European Conference on Artificial Intelligence
2004
Corpus ID: 15945079
This paper presents the clustering algorithm PoBOC (Pole-Based Overlapping Clustering). It has two main characteristics: the…
Expand
Highly Cited
2000
Highly Cited
2000
Construction of speech corpus in moving car environment
Nobuo Kawaguchi
,
S. Matsubara
,
+4 authors
Y. Inagaki
Interspeech
2000
Corpus ID: 6189324
The Center for Integrated Acoustic Information Research (CIAIR) at Nagoya University has been collecting speech corpora in moving…
Expand
Highly Cited
1998
Highly Cited
1998
TELEPHONE SPEECH CORPUS DEVELOPMENT AT CSLU
R. Cole
,
M. Fanty
,
M. Noel
,
T. Lander
1998
Corpus ID: 18082891
This paper describes eight telephone-speech corpora at various stages of development at the Center for Spoken Language…
Expand
Highly Cited
1960
Highly Cited
1960
Fonctions des Corpora allata chez Locusta migratoria (L.)
L. Joly
1960
Corpus ID: 82035844