True to type

0

True to type

I read a report in Le Monde about how the French cryptologist David Naccache managed to discover an intentionally censured (blacked out) word in a CIA print document released by the White House on April 10. The phrase was “operative told an XXXXXXX service…”.

First they used OCR (check out Simson Garfinkel’s recent encomium of the virtues of character recognition technology) to identify the font used, since it determines the number of characters per unit of length (16 mm). Luckily the font – Arial – was proportional rather than monospace, which meant that an ‘i’ letter took up less space than ‘n’. So they used a dictionary to list the possible words (only 1,530!) of 16 mm. Since the target word came after the string ‘an’, this limited the possibilities to 346 nouns and adjectives. Of these, only 7 made possible sense in context (Ukrainian, uninvited, unofficial, incursive, Egyptian, indebted and Ugandan). Given extra-textual circumstances, ‘Egyptian’ was chosen as the most likely candidate.

None of this deciphering was automated, of course, and the actual decoding was hardly earth-shattering. But there’s obviously still semantic mileage in the formal properties of fonts.

Downloadable EU glossaries

0

Tip of the hat to Jost Zetzsche’s very handy Tool Kit newsletter on how to get more out of what I would call language technology. In the latest issue he points to a downloadable European Union multilingual ‘thesaurus’ (a glossary really) called Eurovoc covering all fields in which the European Communities are active. Especially useful for ‘enlargement’ languages.

BTW, thanks to a commentator for the reference to the (decidedly non-downloadable) Eureka terminology server. 

Arabic search engine

0

The Israeli company Melingo has launched Morfix CL, an ‘English-Arabic-English Cross-Language Search with Embedded Translation’ that uses a special engine to handle the massive morphological complexities of Arabic when running word searches. Check this demo

Killer English

1

You’ve probably seen the mass media coverage of the language economics of EU enlargement (€2.55 per Euro citizen a year for Community interpretation and translation). This is one of the rare occasions where ‘language activity’, for want of a better term, is publicly costed in this way, even though the localization industry is confronted by pricing issues every day. One can only hope that people do not confuse that couple of euros per head per year with European translation activity in general.

For a completely different take on the idealized 420 language-pair combination game for Community interpreters, check out the interview that Eurolang, the European minority language news service, published with linguistic rights activist Tove Skutnabb-Kangas on the linguistic power play of European enlargement:

“English is the world’s worst killer language. When a language is learned subtractively, at the cost of the mother tongue, instead of in addition to the mother tongue, it becomes a killer language. So far, most people in Europe have learned English in addition to their own languages, unlike many people in Asia and, especially, Africa, but there are already researchers, civil servants, who know certain things much better in English than in their own languages. Sometimes they are unable to discuss them or write about them in their own languages. This trend may become stronger with the enlargement, also because using English has an even higher status in many of the accession states. In the “best” case, only English and the other official EU languages will develop while minority languages lose out.

“Britain profits hugely from the image that their “brand” of English is the most sought-after. Some researchers have started counting how much the UK and the USA save and/or benefit from others learning their language unilaterally, and suggest compensation. We need to make sure that all languages in Europe are learned, used and developed, in all areas. And of course Catalan, Basque and Welsh should be both official and working languages in EU.”

“This depends on the language policy awareness among the politicians and civil servants. In a comparison in 2001 I looked at which countries had signed and ratified both international and European language and minority related human rights instruments, for instance the European Charter for Regional or Minority Languages and the Framework Convention for the Protection of National Minorities.

“The accession countries were either at the same level or better than the then EU member countries. And they were doing better than the EU countries in several aspects of implementation. There is a lot of hypocrisy in the “old” EU countries which have consistently demanded more of the “new” countries than they have been willing to do themselves.”

“The interpretation and translation costs, even with enlargement, will probably be considerably under 2% of the EU’s administrative budget – this is a very low figure … The costs for supporting linguistic diversity are low, and it is money well spent, according to several economists.

“In the future, when a large part of the world’s population has near-native competence in English, it will not pay off to invest in English only in the way EU schools do now. High levels of English will be like literacy long ago and computer literacy now, something that is a necessity but not a sufficient prerequisite for most jobs, especially well-paid jobs. Competence in other languages will give higher salaries. High levels of multilingualism enhance creativity, cognitive flexibility and divergent thinking – and these are the capital that Europe will need if we want to manage in global competition.

“Europe is linguistically the poorest part of the world, with only 3% of the world’s languages. And we are busy killing off even that small diversity … The critics usually unfortunately know very little about both language policy and economics.”

Balkan language translation training

0

The Balkanika Foundation has launched the first permanent school for translators of all Balkan languages in the Bulgarian town of Sevlievo.

Collateral damage

0

Saddening report on the ˜involvement” of hired Arabic language hands in events in the now infamous Abu Ghraib prison where they are working under sub-contracts for coalition forces in Iraq:

“Titan has been providing translators for the Army in Guantanamo Bay, Cuba; Afghanistan; and Iraq under a long-term contract with the Army Intelligence and Security Command, or INSCOM. Titan will provide as many as 4,800 linguists and be paid up to $657 million, INSCOM spokeswoman Deborah Parker said.

Titan recently reported its revenue under the contract amounted to about $112.1 million in 2003.

INSCOM also awarded a $154.7 million contract in September to CACI “to provide mission support services at INSCOM sites, other national intelligence agency sites, and for other army tactical units worldwide.”

WordPerfect redux

0

WordPerfect Office 12 has just been released by Corel. The original Wordperfect was one of the great milestones for document creators, managers, translators etc back in the late 1980s and early 90s, and was the possibly the first comprehensive word processing package to be thoroughly localized. Corel claims this latest version includes ‘multilingual writing tools for many languages including French, German, Spanish, Dutch and Italian. The suite also includes 40 new templates, the Pocket Oxford Dictionary, 9,500 clipart images, 600 TrueType fonts and more than 180 photos’.

But they didn’t manage global simultaneous release. ‘English versions of WordPerfect Office 12 are available today across North America, the United Kingdom, Australia, New Zealand, and South East Asia. German, French, Spanish and Brazilian-Portuguese versions of WordPerfect Office 12 will be available beginning in the summer of 2004.’

Text analytics

2

Bernard Huber has recently published Textalyser, a web server with a French-English interface that will analyse any input text into statistical data about word tokens. This sort of information is useful for pricing translations, sizing website content and anticipating difficulties in text to speech conversion for a text’s ‘longest’ words.

What text workers need is a fast-reaction tool box including at least word analytics, a concordancer to see words in context, and term extraction capabilities. Ideally accessible by clicking on any word in a text. We should be able to experience electronic words as portals to knowledge about the their lexical dominion, and their given instantiation in a document. But we cannot capture this ‘knowledge’ on a personal hard disc: what might look like a useful add-on to a word processing application actually needs to be web-based to benefit from richer, broader knowledge streams about words and language. Maybe the data that Textalsyer generates about texts, for example, can itself be aggregated to provide a further level of useful statistics about web-wide textual practice. But you probably need some sort of classificatory metadata about the semantic rather than purely formal content of the texts themselves to make this useful.

New life on screen for old languages

0

There seems to be new life for dead languages in movies. The latest beneficiary, of course, is 1st century Aramaic, resurrected in Mel Gibson’s The Passion of the Christ, with help from a California Biblical scholar. Before him Derek Jarman got his actors speaking Latin in the Sebastiane (note the cool vocative), though listeners to Nuntii latini (Finnish radio station broadcasting in Latin) would naturally disagree that Latin was a dead language. I haven’t seen Mel’s film but wonder whether there is an Aramaic version of the title anywhere in the movie.

BBC (and now DVD) buffs might have seen a docu-series by Tony Mitchell called Ancient Egyptians, with actors speaking bits of reconstructed Egyptian. The main story line, though, comes through a spoken English commentary. Opportunities for narratives derived from dead languages (Beowulf, The Epic of Gilgamesh, etc) abound, and lend themselves to the sort of superb computer graphics that morph old landscapes into pristine beauty. And viewers appear to suspend disbelief and accept the pseudo-authenticity of bits of, say, Elvish in Lord of the Rings, just as their elders accepted phoney foreign accents in U.S. movies purporting to tell stories set in Germany, or France. I presume someone can guide us to a website where a tally is being kept on linguistic oddities (not necessarily moribund tongues) in the reel world, whether mumbo-jumbo Egyptian in Mummy movies, or fantasies like Anthony Burgess’s invented ‘prehistoric’ language in Quest for Fire