Skip to search form
Skip to main content
Skip to account menu
Semantic Scholar
Semantic Scholar's Logo
Search 236,658,252 papers from all fields of science
Search
Sign In
Create Free Account
Text normalization
Known as:
N11n
, Normalization
, Text normalisation
Text normalization is the process of transforming text into a single canonical form that it might not have had before. Normalizing text before…
Expand
Wikipedia
(opens in a new tab)
Create Alert
Alert
Related topics
Related topics
8 relations
American and British English spelling differences
Outline of natural language processing
Sed
Separation of concerns
Expand
Broader (1)
Natural language processing
Papers overview
Semantic Scholar uses AI to extract papers important to this topic.
2019
2019
Historical Text Normalization with Delayed Rewards
S. Flachs
,
Marcel Bollmann
,
Anders Søgaard
Annual Meeting of the Association for…
2019
Corpus ID: 196210175
Training neural sequence-to-sequence models with simple token-level log-likelihood is now a standard approach to historical text…
Expand
2018
2018
Text Normalization and Spelling Correction In Kazakh Language
Gaukhar Slamova
,
M. Mukhanova
International Joint Conference on the Analysis of…
2018
Corpus ID: 91185795
. Text normalization is significant step in preprocessing of informal, social media and short texts in the Natural Language…
Expand
2017
2017
On Evaluating the Contribution of Text Normalisation Techniques to Sentiment Analysis on Informal Web 2.0 Texts
Alejandro Mosquera
,
Yoan Gutiérrez-Vázquez
,
P. Moreda
Proces. del Leng. Natural
2017
Corpus ID: 27231913
The writing style used in social media usually contains informal elements that can lower the performance of Natural Language…
Expand
2016
2016
Minimally supervised written-to-spoken text normalization
Axel H. Ng
,
Kyle Gorman
,
R. Sproat
Automatic Speech Recognition & Understanding
2016
Corpus ID: 8625971
Text normalization is the task of converting from a written representation into a representation of how the text is to be spoken…
Expand
2015
2015
DIEGOLab: An Approach for Message-level Sentiment Classification in Twitter
A. Sarker
,
Azadeh Nikfarjam
,
D. Weissenbacher
,
G. Gonzalez-Hernandez
International Workshop on Semantic Evaluation
2015
Corpus ID: 16035780
We present our supervised sentiment classification system which competed in SemEval2015 Task 10B: Sentiment Classification in…
Expand
2015
2015
A Corpus-Based Concatenative Speech Synthesis System for Marathi
Sangramsing Kayte
,
Monica R. Mundada
,
C. Kayte
2015
Corpus ID: 38141627
Speech synthesis is the process of converting written text into machine-generated synthetic speech. Concatenative speech…
Expand
2013
2013
DLSI en Tweet-norm 2013: Normalizaci on de Tweets en Espa~ nol DLSI at Tweet-norm 2013: Normalisation of Spanish Tweets
Alejandro Mosquera
,
DLSI-Universidad de Alicante
,
P. Moreda
2013
Corpus ID: 64457695
on, Tweets Abstract: The lexical richness and its ease of access to large volumes of information converts the Web 2.0 into an…
Expand
2008
2008
UTDallas at TREC 2008 Blog Track
Bin Li
,
F. Liu
,
Yang Liu
Text Retrieval Conference
2008
Corpus ID: 2854723
Abstract : This paper describes our participation in the 2008 TREC Blog track. Our system consists of 3 components: data…
Expand
2008
2008
Bob: A lexicon and pronunciation dictionary generator
V. Wan
,
J. Dines
,
A. Hannani
,
Thomas Hain
IEEE Spoken Language Technology Workshop
2008
Corpus ID: 12537675
This paper presents Bob, a tool for managing lexicons and generating pronunciation dictionaries for automatic speech recognition…
Expand
1993
1993
Pronunciation and text normalisation in applied text-to-speech systems
Margaret Gaved
EUROSPEECH
1993
Corpus ID: 1964184