Automatic capitalisation generation for speech input

نویسندگان

Ji-Hwan Kim

Philip C. Woodland

چکیده

Two different systems are proposed for the task of capitalisation generation. The first system is a slightly modified speech recogniser. In this system, every word in the vocabulary is duplicated: once in a decapitalised form and again in capitalised forms. In addition, the language model is re-trained on mixed case texts. The other system is based on Named Entity (NE) recognition and punctuation generation, since most capitalised words are the first words in sentences or NE words. Both systems are compared when every procedure is fully automated. The system based on NE recognition and punctuation generation shows better results by word error rate, by F-measure and by slot error rat e than the system modified from the speech recogniser. This is because the latter system has a distorted language model and a sparser language model. The detailed performance of the system based on NE recognition and punctuation generation is investigated by including one or more of the following: the reference word sequences, the reference NE classes and the reference punctuation marks. The results show that this system is robust to NE recognition errors. Although most punctuation generation errors cause errors in this capitalisation generation system, the number of errors caused in capitalisation generation does not exceed the number of errors from punctuation generation. In addition, the results demonstrate that the effect of NE recognition errors is independent of the effect of punctuation generation errors for capitalisation generation.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A rule-based named entity recognition system for speech input

In this paper, we propose a rule based (transformation based) named entity recognition system which uses the Brill rule inference approach. To measure its performance, we compare the performance of the rule-based system and IdentiFinder, one of the most successful stochastic systems. In the baseline case (no punctuation and no capitalisation), both systems show almost equal performance. They al...

متن کامل

A Database for Automatic Persian Speech Emotion Recognition: Collection, Processing and Evaluation

Abstract Recent developments in robotics automation have motivated researchers to improve the efficiency of interactive systems by making a natural man-machine interaction. Since speech is the most popular method of communication, recognizing human emotions from speech signal becomes a challenging research topic known as Speech Emotion Recognition (SER). In this study, we propose a Persian em...

متن کامل

Rules for Automatic Grapheme-to-Allophone Transcription in Slovene

The domain of spoken language technologies ranges from speech input and output systems to complex understanding and generation systems, including multi-modal systems of widely differing complexity (such as automatic dictation machines) and multilingual systems (for example, automatic dialogue and translation systems). The definition of standards and evaluationmethodologies for such systems invo...

متن کامل

Robotics Control Using Isolated Word Recognition of Voice Input ( NASA - CR - 155535 ) ROBOTICS CONTROL USING N 78 - 15752 ISOLATED WORD RECOGNITION OF VOICE INPUT

Use of Audio for Robotics Control 1 vii 1. 2. Human Mechanisms for Speech Generation and Recognition 8 2.1 Human Speech Production 8 2.2 Human Speech Recognition 10 3. The Automatic Isolated Word Recognition System 12 3.1 General Description 13 3.2 Feature Extraction 17 3.3 Data Compression and Normalization 27 3.4 Utterance Comparison and Classification 48 3.5 Organization and Operation 65 4. ...

متن کامل

A Comparison of Different Approaches to Automatic Speech Segmentation

We compare different methods for obtaining accurate speech segmentations starting from the corresponding orthography. The complete segmentation process can be decomposed into two basic steps. First, a phonetic transcription is automatically produced with the help of large vocabulary continuous speech recognition (LVCSR). Then, the phonetic information and the speech signal serve as input to a s...

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

Computer Speech & Language

دوره 18 شماره

صفحات -

تاریخ انتشار 2004

Automatic capitalisation generation for speech input

نویسندگان

چکیده

منابع مشابه

A rule-based named entity recognition system for speech input

A Database for Automatic Persian Speech Emotion Recognition: Collection, Processing and Evaluation

Rules for Automatic Grapheme-to-Allophone Transcription in Slovene

Robotics Control Using Isolated Word Recognition of Voice Input ( NASA - CR - 155535 ) ROBOTICS CONTROL USING N 78 - 15752 ISOLATED WORD RECOGNITION OF VOICE INPUT

A Comparison of Different Approaches to Automatic Speech Segmentation

عنوان ژورنال:

اشتراک گذاری