Skip to search form
Skip to main content
Skip to account menu
Semantic Scholar
Semantic Scholar's Logo
Search 237,147,094 papers from all fields of science
Search
Sign In
Create Free Account
Text segmentation
Known as:
Chinese word segmentation
, Word segmentation
, Word splitting
Expand
Text segmentation is the process of dividing written text into meaningful units, such as words, sentences, or topics. The term applies both to mental…
Expand
Wikipedia
(opens in a new tab)
Create Alert
Alert
Related topics
Related topics
15 relations
Cluster analysis
Delimiter
Document classification
Hidden Markov model
Expand
Papers overview
Semantic Scholar uses AI to extract papers important to this topic.
2016
2016
A Character-Aware Encoder for Neural Machine Translation
Zhen Yang
,
Wei Chen
,
Feng Wang
,
Bo Xu
International Conference on Computational…
2016
Corpus ID: 16307400
This article proposes a novel character-aware neural machine translation (NMT) model that views the input sequences as sequences…
Expand
2012
2012
Reducing Approximation and Estimation Errors for Chinese Lexical Processing with Heterogeneous Annotations
Weiwei Sun
,
Xiaojun Wan
Annual Meeting of the Association for…
2012
Corpus ID: 470570
We address the issue of consuming heterogeneous annotation data for Chinese word segmentation and part-of-speech tagging. We…
Expand
2010
2010
Fast-Champollion: A Fast and Robust Sentence Alignment Algorithm
Peng Li
,
Maosong Sun
,
Ping Xue
International Conference on Computational…
2010
Corpus ID: 6734393
Sentence-level aligned parallel texts are important resources for a number of natural language processing (NLP) tasks and…
Expand
2006
2006
Automatic multimedia indexing: combining audio, speech, and visual information to index broadcast news
K. Ohtsuki
,
K. Bessho
,
Y. Matsuo
,
S. Matsunaga
,
Y. Hayashi
IEEE Signal Processing Magazine
2006
Corpus ID: 17586703
This paper describes an indexing system that automatically creates metadata for multimedia broadcast news content by integrating…
Expand
2002
2002
Learning pattern rules for Chinese named entity extraction
Tat-Seng Chua
,
Jimin Liu
AAAI/IAAI
2002
Corpus ID: 5052866
Named entity (NE) extraction in Chinese is very difficult task because of the flexibility in the language structure and…
Expand
Highly Cited
1999
Highly Cited
1999
A Statistical Information Extraction System for Turkish
Gökhan Tür
,
Dilek Z. Hakkani-Tür
,
Kemal Oflazer
1999
Corpus ID: 13429290
Information Extraction (IE) is the process of analyzing natural language text or speech, and collecting information about…
Expand
1997
1997
Document image compression and analysis
O. Kia
1997
Corpus ID: 58763152
Image compression usually considers the minimization of storage space as its main objective. It is desirable, however, to code…
Expand
1997
1997
Lexicon Effects on Chinese Information Retrieval
K. Kwok
Conference on Empirical Methods in Natural…
1997
Corpus ID: 6614339
We investigate the effects of lexicon size and stopwords on Chinese information retrieval using our method of short-word…
Expand
1997
1997
Event tracking and text segmentation via hidden Markov models
J. Yamron
,
I. Carp
,
L. Gillick
,
S. Lowe
,
P. V. Mulbregt
IEEE Workshop on Automatic Speech Recognition and…
1997
Corpus ID: 62561560
We present an approach to the problems of text segmentation and event tracking that makes use of hidden Markov modeling and…
Expand
Highly Cited
1994
Highly Cited
1994
Text segmentation in mixed-mode images
N. Chaddha
,
Rosen Sharma
,
Avneesh Agrawal
,
Anoop Gupta
Proceedings of 28th Asilomar Conference on…
1994
Corpus ID: 58358347
Block based algorithms have found widespread use in image and video compression. However, popular algorithms such as JPEG, which…
Expand