Words that count

0

Sorry I’ve been offline again for a long time. Promise to keep on track now I’m back from la France profonde.

Let’s start back with a spot of fun – at least for English-speakers.

Don’t expect a superior lexeme counter for your translations if you go to WordCount . It’s actually an ‘artistic experiment’ carried out by Jonathan Harris, a designer with Number27. He presents the “86,800 most frequently used English words (from proper nouns to prepositions), ranked in order of commonality” along a line on diminishing font size from left to right. What does he mean by commonality here? Dictionaries suggest it means “The possession, along with another or others, of a certain attribute or set of attributes: a political movement’s commonality of purpose, or A shared feature or attribute”. For WordCount this presumably means that the words share the feature of being “scaled to reflect its frequency relative to the words that precede and follow it”.

What you can do with this artistic toy is either find out how “frequent” a word is (e.g. the last word in the whole list is oddly, conquistador, ahead of items like conflas and fwag), or enter any digit up to 86,800 and see which word has that position in the rankings. Harris takes the data from the 100 million word British National Corpus that covers a very wide range of spoken and written sources. He plans to use WordCount to track “word usage within any desired text, website, and eventually the entire Internet.”

But what people really like about WordCount, apparently, is the fact that the arbitrary word line of relative frequencies sometimes generates surreal phrases: e.g. around translate you get “pussy patting translates geomorphology impasse”. Quite. And around globalization, you find “BASF Hokkaido repugnance globalization Sunderby”. Odd stuff. In other words, we keep on seeing meaning where the machine simply lines up arbitrary forms.

Europe’s Multilingual Companion

1

Most people these days know how to smile at the ‘five year’s away’ promises from the IT sector, especially in the language technology industry. In Europe (and no doubt elsewhere), the language and speech technology community has started building various roadmaps to add a bit more rigor to specifying developmental milestones in such areas as speech recognition and MT. It will be interesting to see whether, like certain other roadmaps, progress follows the vision. Recently, one of the bodies that advises the EU on future R&D programs (the European Information Society Technologies Advisory Group (ISTAG)) published a draft report on the Grand Challenges in the Evolution of the Information Society which tries to crystallize possible futures (circa 2010) into a number of easy-to-picture challenges and solutions.

One of them is devoted to the “serious linguistic challenges” of a very large Union, and proposes “The Multilingual Companion” – “a small portable device making it easy for each person (i.e. in an encounter) to understand the others.” Here’s the vision:

A device of this sort…would rapidly and accurately render translations in text and speech in any other European language desired. Hence, it could be used to take voice dictation, immediately generating the text in any number of languages, thereby streaming the connection between humans and computers. The companion would be extremely useful for disseminating the minutes of meetings to many countries simultaneously, to each in its own language.

The idea behind this challenge is, of course, to focus research in all areas of language technology (speech, learning, cognitive clues for meanings, summarizing, translation, cross-lingual search, etc) towards a concrete practical goal rather than towards a scatter of purely short-term R&D targets, as has often been the case in the past. The model here is clearly the successful VERBMOBIL project run by the German government to produce a mobile real-time translation device for a limited language set.

As the report points out, European IT giants like Philips and France Telecom have been reducing their R&D effort in these areas in recent years, while IBM and Microsoft have boosted research in ‘natural language’ technology. ISTAG’s hope is that the new generation of language technology companies, that in several cases have been spun off from larger organizations, or which emerged from earlier EU language tech projects, will now carry the work forward in Europe to produce the sort of a polyglot wristwatch that will solve our face to face communication problems.

Crystalline

0

Crystal – there’s something about that wonderful lucid, mineral name. Students of linguistics who remember it will be happily amused to see that British jack-of all-language-trades David Crystal has started a company called Crystal Semantics which has just launched a search engine called Textonomy Reveal – “designed to deliver relevant, coherent, and accurate results based on the linguistic sense of the words used in a search.”

“The sense engine which drives Textonomy is the result of a search linguistics development programme which has taken six years to date and involved an investment of over £4 million in lexicographic and encyclopedic research. The dictionary component currently has a coverage of over 200,000 items, and the encyclopedia component contains some 4 million words. The apparatus developed to enable the sense engine to function has received a UK patent, with US patent pending.”

Crystal originally developed his own taxonomy of word meanings for the print version of The Cambridge Encyclopedia, and his work was taken up by the late Dutch company AND Software for various applications. Don’t confuse Crystal Semantics with Translation Crystallization which according to W.A.P.A Translations“>W.A.P.A Translations is a

“new trend… has stealthily been emerging in the translation industry. Instead of companies outsourcing their translation work to many different freelancers all over the world, corporations prefer to have all their work handled by a so-called ‘translation partner’. This partner in translation looks after the clients needs and demands from the beginning of the project until long after the final stage of the ‘crystallization process’ has been completed.”

Back to Textonomy: It seems to work only for English, so there is still time to look forward to a cross/multilingual version along the lines of EuroWordNet. Or will this latest taxonomy/thesaurus effort be gobbled up in the global search engine wars pitching Google against Yahoo and Microsoft? Out with your crystal ball!

80,000 pages of EU law into Irish, anyone?

0

Eurolang reports that the Irish Government is to “initiate a process of discussions with the other EU Member States and the EU Commission with a view to seeking official and working language status for the Irish language in the EU under EEC Regulation 1/1958.”

Pádraig Ó Laighin, who led the Stádas campaign for full working status told Eurolang:

“More people will speak and use Irish in the future as a result of this decision, and the whole raison d’être for learning the language in school, as a first or second language, has been transformed. No longer will Irish speakers be second-class citizens in their own country. There are no guarantees for the future of any language such as Irish, which is under threat, but this decision restores hope and strengthens self-esteem”.

Spivack on metalanguage

0

I don’t usually trust the explanatory power of obvious analogies between the Web and the brain / mind or genes and memes, so I was going to give Nova Spivack’s recent article entitled Minding the Planet: From Semantic Web to Global Mind a miss. But if you want an accessible big-picture approach to the computing (as opposed to the linguistic) concept of metalanguage and its importance for the future of distributed intelligence, try it. Here’s a excerpt on how agents can mark up a text article with semantic data in various different areas of expertise:

We might even imagine that some of these agents are capable of generating new articles and data structures about the original article and linking them together – for example, one agent might generate a synopsis, another might translate it into another language, another might measure the opinions in the article, still another might generate a report based on the conclusions in the article. Because all of this knowledge is expressed using open semantic metadata standards, any program that later encounters any of it can make use of it in its own work, without having to be expressly programmed to do so.

This is already starting to happen in fact – For example, in the blogging community and communities of practice, which in an entirely bottom-up emergent manner, are naturally aggregating, annotating, linking, organizing and prioritizing information. Although there is no central guidance within such knowledge communities, their collective self-organizing behavior results in global information processes that appear to be intelligent. If one were to view the information dynamics of the Web from space – perhaps with a special sensor that could detect and measure these patterns as they emerged – would it not appear similar to the a functional brain imaging scan?

Bloccitanian news

0

Sorry to have been abruptly off-line so long. The problem was part a sudden change in holiday plans, part connection hassles in a village in the S. of France and part keeping an eye on a perfect EU microcosm of mixed-language (FR, DN, GER, SP) kids, aged 5 to 12, who inevitably had English as their lingua franca, with the oldest acting as interpreter between FR and ENG for the FR monolingual Sara and FR-GER bilingual Maha. No Occitan speakers, apparently, though that was the original language of the village in which they were all playing. As it happens, there’s a reference from Eurolang to a new report published on the French site by the International Organization of Francophony (OIF) claiming that “the enlargement of the EU has further weakened the position of the French language”. Nonsense of course — what they are really complaining about is the weakening of the influence of French (geo)political ideas. It is part of France’s genius to identify the content with code. Now back to some real blogging.

Bonn voyage

0

I’m blogging from Bonn, John Le Carre’s famous small town in Germany for a certain generation, where I’m attending a very dynamic second edition of Localization World. The sun is shining, barges slowly plow the Rhine outside the windows of the Beethovenhalle, and the post-event beers are satisfyingly cool. The event is well-attended, well-organized and off to a good start. New term of the day: reverse localization, used by keynote speaker Peter Williamson, Professor of International Studies and Asian Business at Insead France and Singapore. His idea is that not only will a knowledge-driven economy need what we normally think of as localization (adapting products to locale/language) but it will increasingly require the reverse process of ‘localizing’ back to the manufacturer the kind of local(e) knowledge that their developers need about locale sensitivity to design better products. Williamson ended up by challenging the industry to “raise their game”. Meaning presumably that bidirectional localization of the sort he was suggesting will position translation more than ever as the knowledge broker. 

Reverse globalization for Spider-Man

1

Here’s a cute story from the Chicago Sun Times for people who worry about what localization and globalization have in common. Apparently Spider-Man is to be ‘localized’ to ‘India’ (under the name “Indian Spider-Man”), apparently the first comic book hero to undergo such treatment. Some are calling the process “reverse globalization”.

Shakespeare as he woz spoke

0

The UK Telegraph ran a fun story about the London Globe Theatre’s project of putting on a Shakespeare play using the pronunciation as spoken in Will’s time. They’ve got language polymath David Crystal to advise on the sort of English they would have used to perform Romeo and Juliet in 1590.

“It’s more like West Country,” (…) “with a whirred Scots “r” and vowels colored like those in modern French.”

So expect something like ‘weerrforre arrt thu, Rommio’ I suppose. This will no doubt open up another Pandora’s box of ‘authentic’ performance possibilities, similar to using ancient musical instruments, or old amphitheaters. After Mel Gibson’s cod-Latin/Aramaic J.C., we might expect a more linguistically reconstructed Troy one day or even ancient Hebrew or Sanskrit declamations in original vocal garb. Maybe, as multilinguality slowly becomes tamed into the natural modality of a globalized society, ‘language’ itself will take on a more playful, sensual role, like good food or sightseeing, instead of just a survival medium for meanings.

The Shakespeare gig is of course sheer high jinks plus a little intriguing science, but you could also look at is as a ‘de-localized’ version of the play for those who wish to see it – i.e. it intentionally fails to adapt a product to its natural locale. Which brings me to a minor bête noire. In my recent experience of good theater in Europe, there’s still a long way to go before we can deliver appropriate localized versions to unequally endowed audiences.

For example, I saw a wonderful Catalan version in Paris recently of The Threepenny Opera, performed in Catalan because it was an extravaganza stuffed with local references, but the sur-titles designed for an audience who came to a Paris suburban theater were only in English. No doubt the troupe had traveled round Europe with the show and there was no budget to translate for each locale. But is it so hard to do? Same thing a few years ago with a Bob Wilson production in Paris, sung/spoken in German but with English surtitles. No one around me seemed to complain, mind you, possibly because it doesn’t look cool to admit to not mastering the new lingua franca. But now that this sur-title technology is gradually improving (but frankly, why can’t we expect normal text standards with upper and lower case and accents etc, and capable of handling two languages if necessary?), it seems to be a minimum requirement of a continent ostensibly devoted to improving citizens’ rights to be a bit more inventive when it comes to the virtues of theatrical and operatic localization. If all goes well, there should caravans of traveling shows parading their wares from the Urals to the Atlantic in the years ahead. Let’s put some more money into localizing understandable contemporary versions that everyone can enjoy before we blow it all on the great Barrrd!