GermAnProse: 18,000 annotations across four German prose texts
TL;DR
- GermAnProse ships four German short prose texts with more than 18,000 manually created standoff annotations in JSON.
- The paper introduces Characters in Action (ChiA), an annotation scheme covering mentions, speech and character agency.
- The corpus also includes audiobook timing data, capturing pauses between sentences and how long each sentence takes to read.
A new paper in the LaTeCH-CLfL 2026 proceedings introduces GermAnProse, a corpus built around four German short prose texts and more than 18,000 manually created standoff annotations.
The authors describe it plainly. "We present the novel dataset GermAnProse, an annotated corpus consisting of four German short prose texts accompanied by an extensive set of narrative-focused annotations," they write. The annotations cover mentions, speech and character agency under a scheme the paper calls Characters in Action (ChiA), alongside narrativity, semantic verb classes and plot keyness.
One less common addition: reader reception data drawn from audiobook performances, with "timing information for audiobook performances, indicating pauses between sentences and the time taken to read a specific sentence." Everything ships as JSON standoff annotations, released for further exploratory work.
Shared on Bluesky by 1 AI expert
-
Interested in characters, agency, narrative structure or reading performance across complete literary texts? Meet GermAnProse: four German short prose works, fully annotated for computational literary studies. Paper: a…
View on Bluesky →
Originally reported by aclanthology.org
Read the original article →Original headline: Narrative in Short German Prose: A Multi-Phenomenon Dataset for Computational Literary Analysis