A Fact-aligned Corpus of Numerical Expressions

نویسندگان

Sandra Williams

Richard Power

چکیده

We describe a corpus of numerical expressions, developed as part of the NUMGEN project. The corpus contains newspaper articles and scientific papers in which exactly the same numerical facts are presented many times (both within and across texts). Some annotations of numerical facts are original: for example, numbers are automatically classified as round or non-round by an algorithm derived from Jansen and Pollmann (2001); also, numerical hedges such as ‘about’ or ‘a little under’ are marked up and classified semantically using arithmetical relations. Through explicit alignment of phrases describing the same fact, the corpus can support research on the influence of various contextual factors (e.g., document position, intended readership) on the way in which numerical facts are expressed. As an example we present results from an investigation showing that when a fact is mentioned more than once in a text, there is a clear tendency for precision to increase from first to subsequent mentions, and for mathematical level either to remain constant or to increase.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Fact-aligned Corpus of Numerical Expressions Conference Item a Fact-aligned Corpus of Numerical Expressions

متن کامل

The Open University ’ s repository of research publications and other research outputs A fact - aligned corpus of numerical expressions

متن کامل

Language and the Socio-Cultural Worlds of Those Who Use it: A Case of Vague Expressions

The present study is an attempt to investigate the use of vague expressions by intermediate EFL learners. More specifically, the current study focuses on the structures and functions of one of the most common categories of vague language, i.e. general extenders. The data include a 22-hour corpus of English-as-a-foreign-language conversations. A comparison is also made between this corpus and a...

متن کامل

Lexical Bundles in English Abstracts of Research Articles Written by Iranian Scholars: Examples from Humanities

This paper investigates a special type of recurrent expressions, lexical bundles, defined as a sequence of three or more words that co-occur frequently in a particular register (Biber et al., 1999). Considering the importance of this group of multi-word sequences in academic prose, this study explores the forms and syntactic structures of three- and four-word bundles in English abstracts writte...

متن کامل

Annotation of Anaphoric Expressions in an Aligned Bilingual Corpus

This paper discusses a French-English corpus annotated and aligned at anaphoric level. It also presents an annotation scheme based on the study of a detailed corpus featuring different types of correspondences and mismatches. The scheme which is adapted from EAGLES recommendations, supports the alignment at anaphoric level and caters for the different kinds of mismatches.

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره شماره

صفحات -

تاریخ انتشار 2010

A Fact-aligned Corpus of Numerical Expressions

نویسندگان

چکیده

منابع مشابه

A Fact-aligned Corpus of Numerical Expressions Conference Item a Fact-aligned Corpus of Numerical Expressions

The Open University ’ s repository of research publications and other research outputs A fact - aligned corpus of numerical expressions

Language and the Socio-Cultural Worlds of Those Who Use it: A Case of Vague Expressions

Lexical Bundles in English Abstracts of Research Articles Written by Iranian Scholars: Examples from Humanities

Annotation of Anaphoric Expressions in an Aligned Bilingual Corpus

عنوان ژورنال:

اشتراک گذاری