While the U.S. Government is recruiting madly to translate and teach Arabic or develop Arabic language automation systems, an EC funded project – NEMLAR – has just produced a handy report on available digital speech and language resources and analysis tools in European and Mediterranean countries (plus U.S. Resources from the Linguistic Data Consortium), based on a survey of key academic and commercial players.
Section 4 gives a list of what experts consider to be their desiderata for resources and tools, whose length suggests Arabic will not be engineered into our language infrastructure very soon.
It’s also hard to know from EU security authorities (EUROPOL and the 25 or so National Intelligence units) how strategic Arabic is actually considered in the wake of the recent geopolitical events, and therefore how much they would fork out to invest in scalable, operational search and translate systems.
The French defense establishment, for one, has long been researching into Arabic language technology, but I notice it recently contracted the Canadian company Alis to provide some translation gisting support. But this may tell us more about the contrast in PR styles between North America and Europe than actual contracts for developing Arabic language technology.
New report on Arabic digital language resources
Winning the information war
Very interesting article by Samuel G. Freedman in the New York Times on the inadequacy of U.S. government attitudes to strategic language study in an age of global communications. Some quotes:
“Over more than three decades, as the support for language study was written into other federal laws, a steady stream of 30,000 or more American university students took Russian courses each year.(…)
“Meanwhile, of more than 1.8 million graduates of American colleges and universities in 2003, exactly 22 took degrees in Arabic, according to Department of Education statistics. (…)
“Richard Brecht, a former Air Force cryptographer who is executive director of the Center for the Advanced Study of Language, a joint project of the Defense Department and the University of Maryland based in College Park. “Five billion dollars for an F-22 will not help us in the battle against terrorism. Language that helps us understand why they’re trying to harm us will.”
Pictographic writing
For something completely different, check out the only ‘living’ pictographic language (think Mayan or Egyptian hieroglyphs) in the world today used for ritual purposes by the Naxi people of Yunnan Province in China. There’s a rather academic online exhibition of some of their manuscripts at this Library of Congress site.
U.S. Language Map
Lots of blogs are circulating the good news about the Modern Language Association’s MLA Language Map, which “displays the locations and numbers of speakers of the thirty languages most commonly spoken in the United States.” Just the ticket for a multicultural marketer trying localize a product or service down to that single last customer who speaks Hungarian in North Dakota. Site usability for this resource leaves a lot to be desired, since large maps are inherently difficult to manage on small screens, and in any case it is designed as an educational tool. But it sets standards that other organizations might well apply to language-mapping other geographies in a visually appealing, regularly updated and accurate way. If you happen to notice the distinct lack of Amerindian languages on the MLA map, there’s a good resource site for this Sprachbund here (thanks to André Cramblit for the reference). BTW, if you’re in search of the latest traditional facts about ‘countries’ and their ‘languages’, the 2004 edition of the CIA World Fact Book has just gone live.
More vendor newsletters
Received new newsletters from two young language tech companies this week. Newsletters are obviously a well-established media for keeping in touch with clients and prospects, but for very small companies they carry an inevitable cost. So it’s good to hear from acrolinx, the German supplier of acrocheck (which helps control consistency in source documentation due for localization), and also from Lexicool, a French provider of text analysers, and various terminology resources. I’d be interested to known if anyone has kept a tally of regular customer-bound media such as these published in the language industry and localization world. My suspicion is that as more tools (RSS etc) come into play for aggregating a plethora of scattered news sources into relatively manageable dashboards, firms may turn to competitive blogging as way to gain customer mindspace. But the newsletter is still a highly brandable media.
Translator slavery in India!
In an interesting article in India’s leading tech e-news paper, Venkatesh Hariharan, co-founder of Indlinux.org, talks about the speed at which Indian software is localizing,
“We are all set to launch our very first Gnome CD that supports most of the major Indian languages in the next couple of months. The MOSIC based bootable CD ROM has been christened ‘Rangoli’ – which literally means an array of different colours. Just a few months back, Gujarati and Punjabi were no where on the local language radar. With the acceptance of Linux desktops growing steadily, and its maturity now proven, interns were hired from Gujarat for translation – an initiative taken by a group called Utkarsh. To give an example of the commitment of project volunteers, a translator actually sat for 19 hours at a stretch to convert 2000 strings of Linux into Punjabi!â€â€
Punjabi, Marathi, Bengali, Malayalam, Hindi and Gujarati, among others, are supported in this freeware version of Gnome.
Belarusian spell-checker needed for English
The useful DMEurope news service alerted me to the fact that the first Belarusian spell-checking programme for MS Word 2000, (Litara 1.0), has been developed by Fedor Piskunov and Professor Gennadi Tsichun from Minsk State Linguistic University. Claiming to be the largest electronic dictionary of Belarusian in the world, it contains some 170,000 entries and 2 million word-forms in the two variants of Belarusian spelling – official (modern) and traditional (classical) – and has an additional option of translating from the modern spelling to the classical one. Download it free of charge from here. When it comes to statistics, I was sadly amused to see that I managed to google 117,000 links to Belarussian, and a more promising 489,000 to Belarusian.
Newsletter on MT implementation
CrossLanguage, possibly the only company in the world dedicated to providing machine translation implementation services (as opposed to actually offering MT technology), has issued its first newsletter (sign on at the site). In it, Jaap van der Meer introduces the concept of transformer companies:
In the transformer age, customers will need translation at an ever faster rate, sometimes even in real-time. Volumes keep growing, yet there is still tremendous pressure on costs. The old subjectivity that underlies quality evaluation is rapidly disappearing. Translation quality will be measured more easily using standard metrics. The emphasis, meanwhile, will focus on accuracy, utility, and consistency, rather than ‘style’. Finally, an industry that has been inspired over the years by a global projection model, in which English content is translated into forty or more languages, needs to transform itself, and suddenly learn to supply the same language pairs – but in reverse order.
Visualizing language
Visual Thesaurus is a fun, for-sale tool that represents synonymy (and antonymy) relations between English words as graph displays looking like Calder mobiles of interlinked word nodes. I haven’t examined it in detail, but you can click on each related word within a synonym graph to generate new graphs around your term, and also call up boxes with definitions and examples of usage for each term. The actual length of the line linking a word to a synonym is probably meaningful in itself (the longer the line the more ‘distant’ the semantic relationship with the core word or something?).
Fascinating though such a dynamic presentation is, I don’t personally find this sort of interface to language knowledge any more useful in practice than the standard listing of terms in a text-based thesaurus. It’s attractive to the eye, but what can you really do apart from click and watch? It doesn’t link to your documents, or drive a search engine. But its great virtue is to give a momentary illusion of language as a dynamic system of virtual relationships, not just a string of signs on a page. And this might open up interesting new directions for dynamically representing the meaning and style of texts themselves and their relations with each other in ‘knowledge space’. So here’s a wee riff on what you could do with better visualization tech.
Historically, techniques available for ‘representing’ language (e.g. to teach it, or show how two languages or family of languages compare, or demonstrate etymology or syntax etc) has boiled down to leveraging the affordances of:
. print (e.g. using special characters such as the IPA alphabet, and word lists and tables), and
. diagrams – trees to show syntactical structure or historical relations between languages, or transition networks to show relationships that underlie structures.
Phonetics (and now speech technology for pronunciation training) have benefited from spectrograms that show formant behavior for given pronunciations, but they hardly speak to the naïve inquirer. No one seems yet to have invented interesting digital visual models of a given natural language in its dynamic entirety, or of how humans might represent their linguistic knowledge.
Yet visualization techniques for displaying various sorts of knowledge are now going mainstream. You have search engines such as Mooter, Grokker, or The Brain that show search results as shimmering globes of categories or networks of nodes instead of text lists, or enterprise knowledge management tools that present content in terms of visual ‘maps’. Try this from xrefer’s new Research Mapper for a search on ‘machine translation’. There’s even a site that displays looping links between groups and singers according to the query you type in, and masses of other attempts to to use geometry to replace text as the inherent interface to web links.
Now that there are reports of 3D screens about to reach the desktop, isn’t it time for someone to develop a few tools to expand the dynamic display space for language and textuality in general? They might not serve any pressing need for information collection or business process optimization, but they might play a role in teaching anything from writing to linguistics. I’m think of such visual applications as:
. showing dynamic ‘films’ of alternative parses of ambiguous sentences,
. demonstrating phonetic change over time in a fun way (e.g. the Great Vowel shift in English),
. tracing etymologies as visual journeys through time, space and semantic fields
. showing the translation process from the inside: animating how choices are made, how a rule based or a data-driven MT system does its job,
. showing how a speech recognition device uses a database to map incoming acoustic entities onto language models and optimizes its choices,
. displaying ‘complex’ morphological processes in languages such as Russian or Arabic,
. and eventually, showing how a total stretch of speech can reveal its composition through the massive interaction of different types of linguistic phenomena, from prosody through syntax to semantics.
I would imagine being able to click on any link in any part of these displays to enter further into a particular dimension of the model in question, and in this way explore language itself as a system of systems. This kind of application of visualization technology might start as a teaching device, but there’s always someone out there who would hijack it and use it for something more adventurous such as a video game or even an ad. You could even offer it to translation clients who want to watch what happens to their documents rather like a FedEx user tracking their parcel through the geographical maze of the delivery process.