cdh.princeton.edu web signal

Princeton publishes Cairo Geniza data essay for AI researchers

TL;DR

  • Rebecca Sutton Koeser and Marina Rustow published a data essay in Journal of Cultural Analytics 2026, 11(3):1–35, cataloging the Princeton Geniza Project.
  • The released dataset covers 35,855 documents, 1,802 people, and 486 places from the Cairo Geniza (c. 950–1250), archived on Zenodo and updated quarterly.
  • The essay documents modeling decisions, including the many-to-many relationship between physical fragments and the documents written on them.

The Princeton Geniza Project has published a formal data essay describing its corpus of 35,855 documents, 1,802 people records, and 486 places drawn from the Cairo Geniza, the trove of medieval material (c. 950–1250) preserved in the Ben Ezra Synagogue in Fustat. The essay, by Rebecca Sutton Koeser of the Center for Digital Humanities and project PI Marina Rustow of Near Eastern Studies, appears in the Journal of Cultural Analytics (2026, 11(3): 1–35) under CC-BY 4.0, with the dataset itself archived on Zenodo and refreshed quarterly.

"Providing both a well-structured dataset, as well as the contextual information of its creation, enables a broader community of researchers to engage with the Geniza material," Koeser and Rustow write. The essay walks through the modeling choices behind the corpus, including the many-to-many relationship between physical fragments and the documents written on them, a distinction that matters when a single leaf carries more than one text or a single text is scattered across leaves in different libraries.

The Princeton announcement positions the release alongside OpenITI for Arabic texts and KTIV for Hebrew manuscripts, framing the corpus as one meant to be "accessible for data science and AI researchers and historians alike." The scope of downstream tasks — HTR, entity linking, network history — is left to whoever picks it up.

Shared on Bluesky by 2 AI experts