학술논문

The Natural Stories Corpus

Document Type

Working Paper

Author

Futrell, Richard; Gibson, Edward; Tily, Hal; Blank, Idan; Vishnevetsky, Anastasia; Piantadosi, Steven T.; Fedorenko, Evelina

Source

Subject

Computer Science - Computation and Language

Language

Abstract

It is now a common practice to compare models of human language processing by predicting participant reactions (such as reading times) to corpora consisting of rich naturalistic linguistic materials. However, many of the corpora used in these studies are based on naturalistic text and thus do not contain many of the low-frequency syntactic constructions that are often required to distinguish processing theories. Here we describe a new corpus consisting of English texts edited to contain many low-frequency syntactic constructions while still sounding fluent to native speakers. The corpus is annotated with hand-corrected parse trees and includes self-paced reading time data. Here we give an overview of the content of the corpus and release the data.

Online Access

Open Access (Arxiv) Find it@PNU

이메일

부산대학교 도서관

Online Access

메일 발송