The 40 language dash

0

Just before Christmas last year, a French computer geek called Alex Lemaire managed to break the record for the mental arithmetical calculation of the 13th root of a 100 digit number. He did it in 3.62 seconds. He’s now set another global challenge to like-minded calculators – find the 13th root of a 200 digit number (Libération last week) in the shortest possible time.

More interestingly, this memorization hobbyist is also trying to learn 40 (modern) languages simultaneously. By that, I guess he means languages with Latin alphabets, and “learning” probably means memorizing lists of facts about or sayings in a language, rather than demonstrating natural conversational fluency. But there are as yet no details about what he’s doing.

I can’t find a Guinness Book of Records type performance benchmark for this sort of skill, but have the feeling that language hobbyism is set to grow. This could surf on the growing interest in inventing whole languages, and could extend to competitions such as memorizing word lists, speed translating and so on, rather as speed typing was a competitive sport in the early 20th century. As we shift from an era of information scarcity to one of web-driven glut, our sense of ‘language knowledge’ may well evolve into something far more entertaining and competitive, as tongues clash in cyberspace.

Discussions about record-breaking polyglossia, however, have a venerable half life on sites and discussion lists all over the web, even though the notions of “speaking” or “knowing” are notoriously slippery. The most cited (Western) lingo champ is Cardinal Giuseppe Mezzofanti who was reputed to speak thirty-eight languages perfectly. But I was intrigued to come across Plutarch’s reference to Cleopatra’s fairly extensive multilingualism in Nicholas Ostler’s constantly interesting book referred to in a previous post. According to Plutarch in (Thomas North’s translation), Cleo’s guests found that:

It was a pleasure merely to hear the sound of her voice, with which, like an instrument of many strings, she could pass from one language to another; so that there were few of the barbarian nations that she answered by an interpreter; to most of them she spoke herself, as to the Ethiopians, Troglodytes, Hebrews, Arabians, Syrians, Medes, Parthians, and many others, whose language she had learnt; which was all the more surprising because most of the kings, her predecessors, scarcely gave themselves the trouble to acquire the Egyptian tongue, and several of them quite abandoned the Macedonian.

Presumably the Troglodytes spoke a version Berber as spoken in Libya and eastern Numidia (today’s southern Tunisia) where there still people who escape the midday sun by living in caves.  She would have conversed with Mark Antony in Greek (presumably her mothertongue, coming as she did from a Macedonian background). Her story raises the intriguing subject of the glosserotics of multilingual love affairs in history. Interestingly the expression in either singular or plural cannot be found on Google as yet.

Longtemps je me suis…

0

Why can’t Americans read the last volumes of the updated Penguin translation of Proust’s A la recherche…? Read this article from Slate on copyright madness at work.

The dyslexicographer

0

Margaret Marks at Translation Blawg rightly wonders what on earth the Webster’s Online Dictionary (WOD) is all about. Although there is quite a lot of background information available on the site, I decided to find out from its creator Phil Parker. Here’s the score.

A Professor at Marketing at INSEAD, the European business school, Philip Parker was born dyslexic. This meant he found reading dictionaries – lists of words and their definitions – much easier than sustained prose, which demanded too much time to decipher. So over the past 30 years he has been collecting dictionaries of all kinds. Around the year 2000, large dictionaries on the web started charging for ‘premium’ words of the sort he needed in his research and that really “pissed him off”. So he decided to leverage the definitions he had collected from his own store, borrowed the out-of-copyright ‘Webster’ badge, and started building WOD, which he intends to make the biggest multilingual dictionary site on the web.

He was lucky since he had loads of help from academic and other assistants, benefited from donations of out of print dictionaries and word lists, and was able to finance the whole thing himself. He even uses a firm in Togo to keyboard in content. This summer he hopes to upgrade the site to feature dictionaries covering 600 languages (10% of the world’s current language population), and in the case of existing site languages such as Spanish, he hopes to increase the entry count from around 100,000 to 600,000 entries.

To give global coverage he is working in a sequences of passes. The first pass was to work by time zones, taking a location such as Europe and collecting dictionary materials for all ‘major’ languages. The second pass, now under way, is to include ‘secondary’ languages (say Maltese in Europe). Next year, he plans to start the third pass by incorporating locally endangered languages, using volunteer help where necessary. One technique is this: he donates a computer and a small stipend to missionary children (e.g. for Tarahumara in Mexico) who then create a local language/English dictionary.

What’s next, once he’s got all these bilingual word lists? Create a total lexical linker, whereby you can click from any word to its equivalent in any other language, using English as the underlying pivot language. An “N-dimensional cube of words in every language to every language,” as he puts it, that will by this summer be the world’s largest compilation of language items ever produced. His content currently weighs in at around one terabyte.

How useful is Phil’s site proving? He reckons it is among the top ten sites used to search Arabic words in Arabic script, since the whole hoard has been programmed for Unicode. And because the Webster word is a synonym for ‘dictionary’ for Americans (as Kodak once was for cameras or Google is for search engines) WOD ranks between 5 and 7 on, well, Google for ‘Webster’ out of about 150 ‘Webster’ sites on the web these days. Probably the best way to appreciate the ambition of Phil Parker’s site is to search the term Webster itself, and see the degree of encyclopedic potential – words, images, statistical findings from corpora, sign language versions, et alia multa – that he is trying to pack into what he calls a hobby. But the definitions don’t include a more recent decomposition – web + ster (as in napster) – a linguistic peer to peer resource.

Help not pity for the poor immigrant

0

Remember Bob Dylan’s song on John Wesley Harding:

I pity the poor immigrant

Who wishes he would’ve stayed home

One way to help immigrants feel less deterritorialized is to follow this example from Norway, and publish online dictionaries expressly designed for this constituency. This particular set offers Norwegian-Tamil and two dialects of Kurdish. Now that immigrants may well be computerate and use internet cafés to keep in touch with people back home, providing easily updateable online resources makes a llot of sense. Maybe this is a Scandinavian specialty; I notice the The Swedish Schoolnet (funded by the Swedish National Agency for School Improvement offers a Swedish-Bosnian and Swedish-Croatian word hoard as well. 

Localized English

1

Newsweek Europe (probably need to subscribe) this week has a feature on Global English. The basic argument is that a) English as a skill is perceived as the key to jobs and a future, and b) English as a speaking system is simultaneously fragmenting into (national) local Englishes congealing around regional intermediate standards (Indian, Caribbean, African, Asian etc). This more or less echoes what David Crystal has been saying for some time.

In the article, Crystal is quoted as saying that never before in history has a population of second language speakers outpaced native language speakers, as is the case with English today, when you stack up Indian, African and other regional Englishes against the US/UK/Oz/NZ heartland. This may not be true. What about Latin in say the 4th century? Speakers in an arc from Palestine up to modern Romania, across Germania to the Atlantic coast of Gaul and down to Iberia might well have outnumbered pedigree Romani. The best place to find out more about all this in Nick Ostler’s splendid new book Empires of the Word, which offers a highly erudite yet stylish history of the world through its great languages.

Here’s another way to read the English story: the mindware product known as the English language is in the process of being informally localized. This is not due to an intentional push process engineered by the supplier, but to a pull side-effect created by local demand. The result is a hybrid mix of English and local language expressions that meet various communicational needs in the local community. The street is hijacking the language and retooling it with local smarts. Obviously school kids and others strive hard to gain a ‘correct’ core English accent when they speak, but we all know that you end up with a semi-localized accent unless you are extremely gifted, or young enough or brought up as a child in more than one linguistic culture. Do I speak French with a bit of an English accent? No, I speak localized French.

What’s best?

1

Is there a ‘best’ automatic translation system? Sure. But you can only know by constantly testing them all across all languages against a vast range of specific tasks. In other words, never. “Best” looks like a Platonic idea, not a statement of fact.

Lots of blogs and sites these days are quoting the Language Weaver tag line – “the best translation systems in the world”. Does this help or hinder perceptions of translation automation, given the highly relative quality of any translation? Or should L W have taken a leaf out of Heineken’s book and added a modifier. The beer advertisement we sometimes see (at least in Europe) has “probably the best beer in the world”. Again, meaningless, but carefully nuanced to engender that vital suspension of disbelief we need when reading ads.

As it happens, Language Weaver has a best rival – an outfit called Delta, which appears to market a Brazilian/Spanish-English system for 125 bucks a shot, a far cry from what L W’s Arabic to English system would cost. Yet when you check to see how many others are trying to best the market, the results are disappointing. According to this Google search, “best machine translation systems” – note the plural – scored 65 hits, and “best translation system” managed 145, most of them from Delta. With a hundred or so citations out of a billion web pages, it clearly doesn’t matter to most people whether any systems are “best” or not.

That plural in Language Weaver is, however, an astute move. Unlike Delta (and many others) the company claims to develop language-specific systems (Arabic, Chinese, Hindi, French, Spanish to English), based on the analysis of statistical patterns of linguistic type phenomena in large bitexts. In other words, it has some sort of “technology” (recently reported on here probably causing the bloggorhea attack) that it deploys in “systems”.

Most of us, however, have tended to think in terms of the X translation system – that is, the underlying engine in X, that somehow has linguistic knowledge hand-coded into it (a dictionary of say 100,000 terms and 3,000 grammar and morphology rules) to drive whatever language pairs are on offer. Perhaps by going superlative about its nonce systems, rather than its core “machinery”, L W is shifting the focus from technology to task diversification. And in the L W case, translation tasks appear to be addressed by leveraging the deciphered output from previous similar tasks.

Let’s hope that a provocative “best” attracts competitors to this space. Plenty of room for more “systems”, even if there are only one or two effective core technologies.

Collective translation of the Encyclopédie

0

The Encyclopédie Diderot et d’ Alembert is often cited as an early (1750s) beacon of choral authorship (140 contributors) in a traditional European culture largely featuring solo acts. An online version has been available here for some time for paying subscribers only, under the auspices of the French Centre National de la Recherche Scientifique, and the Division of the Humanities, the Division of the Social Sciences, and Electronic Text Services of the University of Chicago.

Now an English version is being stitched together at the University of Chicago as part of the ARTFL project. The translations are, like the original, being carried out as a collective effort. Maybe amateur is the more appropriate adjective:

You don’t have to be an expert translator to contribute to this project, but because it is a collaboration among volunteers, we do not have the resources to edit the translations we receive. We therefore count on contributors to have sufficient command of French, English, and the topic to produce an accurate translation in readable, correct English. Translations can be submitted in either MS-Word or Word Perfect and do not require any special coding.

So far they’ve only managed to translate about 450 of the 72,000 articles that made up the original Encyclopédie. The total base contains 20.8 million tokens (words) or 400,000 types (unique forms), so the average length of an article is only about 300 words. After human translators have completed say 2,000 sizeable articles (c. 600,000 words) maybe someone might be able to prime a translation automation system to manage large chunks of the remaining 70,000 articles. And give a network of students around the world extra credits for post-editing the output.

One nice touch on the website is a facsimile translation of the famous Tree of Knowledge, an attempt to develop an 18th century organon of how knowledge fits together – usually in satisfying triads of categories. But alas, it does not seem to include translating as a fruit hanging from a twig on the ascending Arts of Communicating/Logic/Science of Man/Philosophy/Reason branch.

Here’s Lingster

0

Seems to me there is a knowledge hole in the language technology communication space. There are lots of R&D sites and community events, a few dedicated news sites such as here and here that relay press releases, and one notable publication that is gracious enough to host this blog. What’s missing is any sense of what gets language technologists going. So here’s the first of an occasional series of quickie chats with some lang tech rockers and rollers to find out what’s cooking – and why.

Lingster is an open source initiative to collect and share dictionaries for innovative translation automation systems. It was started by Olga Beregovaya, who works by day as a computational linguist for a large software developer in California, and whose night job is building a multilingual content management system with computer-assisted translation capabilities. Olga is a graduate of St. Petersburg State University, and UC Berkeley.

“I realized I needed lots of lexicons and glossary data. What was available consisted of very plain monolingual word lists, which makes multilingual alignment hard. And they weren’t marked up for morphology or part of speech.

“You can find material on the web but if it’s free and available it’s crap. If it’s decent you have to pay US$ 15,000 or so which is too much for me. I am planning to include 15 language pairs in my application, and I need up to date terms and expressions. So I started lingster as an open source dictionary portal so others could help me and I in turn could help them. 

“The idea is to identify and contribute linguistic data in the form of lexicons and glossaries, idiom lists etc. The next step is to add in part of speech and grammatical information to expressions, and then the R&D community can download the resulting resources for free. This way we all benefit.

“A lot of other people seem to be suffering from the data problem, as the interest in the last few months has been terrific. Lots of linguists see the value of this sort of initiative. Especially since web crawlers programmed to harvest dictionary data end up by hitting Intellectual Property problems. 

“Quality is obviously an issue. We have a filtering system to validate terms using community voting, with a volunteer moderator for each language who throws out unacceptable equivalents. We use a UTF8 control mechanism to stops garbage characters getting into the files, and a conversion mechanism for file formats.

“There is no charge for the software which is all open source, with no hidden dictionary data. We welcome the concept of a community of development, since contributors will be rewarded according to the open source licensing system. Later on, we shall set up a subscription system for those wanting truly premium data.”

And what exactly is Olga’s pet project? An elearning facility that turns any web resource into a potential language learning resource, with instant word lookup. The beta’s due out in a couple of months.

20,000 words under the C:

0

It’s Jules Verne year, the centenary of his death. Blogos has mixed feeings about celebrating the demise of this garrulous if entertaining writer.

The world of multilingual technology fixes has been badly treated by futurists of all kinds. All I could dredge up from Jules V was a banal echo in the nearly unreadable Tribulations of a Chinaman in China of Edison’s hope that recordable wax cylinders would replace handwritten letters. Edison was sort of right (voice mail), but wrong by a platform or two (he thought in terms of sending cylinders by escargot mail).

In hommage, how about this ingenious effort by Gnoetry, a statistical text analyser (+ human minder) that generates human/computer poetry from out-of-copyright texts in rank prose?

These twelve-line blank verse poems were composed with Gnoetry 0.2 and are based upon the statistical analysis of Jules Verne’s 20,000 Leagues Under the Sea. The 0.2 interface allows the human collaborator to make choices by regenerating text on the word, phrase, sentence, and stanza levels; however, no post-composition edits were made by the human author, except where he found the words “continued” and “replied” as arguments for the insertion of quotation marks, capitalized proper names and italicized one foreign word.

A quick footnote. Looks as if the Gnoets are on the way to automating the cadavres exquis model of text production by generating poems out of a combination of two or more original texts.That silence-to-say-goodbye writer William Burroughs used to suggest that by interweaving (i.e. cutting up or folding in printed texts) texts from two different authors (Rimbaud and Newsweek would have been his sort of choice), you could generate a ‘third mind’, a form of random truth that emerges from such forced textual congress.