Dan Cohen

Enhancing Historical Research With Text-Mining and Analysis Tools

Open BookI’m delighted to announce that beginning this summer the Center for History and New Media will undertake a major two-year study of the potential of text-mining tools for historical (and by extension, humanities) scholarship. The project, entitled “Scholarship in the Age of Abundance: Enhancing Historical Research With Text-Mining and Analysis Tools,” has just received generous funding from the National Endowment for the Humanities.

In the last decade the library community and other providers of digital collections have created an incredibly rich digital archive of historical and cultural materials. Yet most scholars have not yet figured out ways to take full advantage of the digitized riches suddenly available on their computers. Indeed, the abundance of digital documents has actually exacerbated the problems of some researchers who now find themselves overwhelmed by the sheer quantity of available material. Meanwhile, some of the most profound insights lurking in these digital corpora remain locked up.

For some time computer scientists have been pursuing text mining as a solution to the problem of abundance, and there have even been a few attempts at bringing text-mining tools to the humanities (such as the MONK project). Yet there is not as much research as one might hope on what non-technically savvy scholars (especially historians) might actually want and use in their research, and how we might integrate sophisticated text analysis into the workflow of these scholars.

We will first conduct a survey of historians to examine closely their use of digital resources and prospect for particularly helpful uses of digital technology. We will then explore three main areas where text mining might help in the research process: locating documents of interest in the sea of texts online; extracting and synthesizing information from these texts; and analyzing large-scale patterns across these texts. A focus group of historians will be used to assess the efficacy of different methods of text mining and analysis in real-world research situations in order to offer recommendations, and even some tools, for the most promising approaches.

In addition to other forms of dissemination, I will of course provide project updates in this space.

[Image credit: Matt Wright]

4 Responses to “Enhancing Historical Research With Text-Mining and Analysis Tools”

  1. PhDinHistory said on February 5th, 2008 at 1:32 am

    That’s awesome. I had hoped this day would come. If you can develop software that can understand texts and find meaning in them, as opposed to just extracting and manipulating their information, I think you will really be onto something. But that may require revisiting our debate about the semantic web.

  2. Steve Ramsay said on February 5th, 2008 at 2:13 pm

    Congratulations, on behalf of the MONK Project! The more the merrier, we say. ;)

  3. Heather Munro Prescott said on February 6th, 2008 at 8:30 pm

    This sounds fascinating. Sign me up!

  4. schrattenkalk.com said on March 15th, 2008 at 4:26 pm

    [...] for history (as a subject) and historians in general. Also on history, but more technical, an article by Dan Cohen on the research on tools for researchers. There is not only blogs, but also podcasts. [...]

Leave a Reply