Unsworth,
John. "New Methods for Humanities Research." The 2005 Lyman
Award Lecture. National Humanities Center. Research Triangle Park, NC.
11 Nov. 2005. <http://www3.isrl.uiuc.edu/ ~unsworth/lyman.htm>.
Tag: text mining
Text to Map
Two
recent
entries drew my attention to the Gutenkarte
project, a series of scripts and processes that renders place-names
appearing in a given text and locates them on
a map. The Gutenkarte site announces
future plans for the project, including a wiki-like annotation add-on that will
enable a group of users to collaborate in expanding the place-name information
and related contextual relevance (one day to include digital images and video?).
The project bears many similarities to Franco Moretti’s survey of the shifting
geographies of village life in the nineteenth century. Moretti’s analysis often
moves beyond standard place-names to include positions of and distances between
people and things known to be in particular places. These he distinguishes as
geometries; plotted, they are more like diagrams than maps, he tells us (54).
The Gutenkarte project is not yet as refined as Moretti’s work; mining a text
for toponyms depends on the database’s tolerance place-name ambiguity and
spelling variations (among other things I probably don’t understand). Still,
despite the obvious limitations, the motives underlying Gutenkarte present an
affirmative answer to one of Moretti’s guiding questions, "Do maps add anything,
to our knowledge of literature?" (35), even if it is being applied to literary
texts from the Gutenberg Project for now.
Unitization Reports
On a break from writing end-of-semester papers for CCR651
and GEO781, I thought I’d shock each of them into a list of noun and noun
phrases by applying the same methods we’ve strung together for CCC
Online. Et voila! The lists aren’t meaningful in quite the way a
sentence-long summary would be. Yet that’s the point. They’re
differently meaningful, suggestive. Maybe even generative if I can trace
through some of the terminal knots tomorrow.
Disappoint-ensity
When a watchful mentor emailed me a link to
Attensity over the
weekend, I was encouraged, finding, from a quick glance at their web site, that
some of the same data-mining they market, as their hallmark, matches up with a few
of the configurations defining a project I have underway.
Attensity claims to process data and analyze that which is otherwise difficult
to discern. Their software churns away like some kind of high-powered
heuristic meaning-cruncher–a processor of mass quantities of text into readable
metadata.
NORA (non-obvious relationship awareness) for large-scale discourse.
So I sent them an email inquiring about the whole plot,
all-the-while-recalling that the only other text parser I looked up asked $2k
per year for software licensing. Did I mention I’m a grad student?
Well, I did mention that in the email to Attensity, the email inquiring about
their project. And I heard back today–a polite note, something about
serving the US Intelligence community and a starting fee of $50,000. And
something about wishing me the best in my quest. Thanks. But no
thank-you.
The
proliferation of textual analysis apps self-identifying as the devices built
to root out terrorism piques my interest. Perhaps because of the sheer
volume of text to be analyzed for particular patterns (suspicious patterns!
watch the parentheticals, Attensity!), mass-discourse text parsers are up and
coming. And so unbelievably over-priced that they’re no use whatsoever to
the project I’m working on.