Proper Names

0

The Adamic power to name things has gone novaburst in the last 50 odd years since people all over the world started hitting computer keyboards and experiencing popular culture via mass media. Above all, this is true of proper names, those which point to but do not define the intension of objects in the world. Just think of the myriad of file names, product names, user names, passwords, rock groups, website domain names, pseudonyms, names of companies, and nonce names of all kinds across all languages that have streamed through our lives. Names for people are an interesting area for comparison, even though the issue there is using the ‘naming system’ rather than necessarily new inventing names. New media activities have swollen the number of discrete ‘names’ by orders of magnitude, I would imagine, when compared to the previous half century. If you can’t think of a new name, you can even get software to invent brand new names.

Naming new products is always a risky area for the multilingually challenged and there have been plenty of silly stories about products shooting themselves in the foot due to insufficient preparation. The reason is that cognitively speaking, native speakers tend to block out the ‘intensions’ of name words in their language(s), while foreign speakers often perceive the component meanings where they don’t resonate for the native. Just think of ‘bush’ or ‘gold man sax’ or ‘brad pit’ or ‘proctor & gamble’ or ‘thatcher’, ‘potter’ or ‘king’.

If you are wondering whether Japanese names have any such hidden wealth for foreigners, check out this site which attempts to unpack meanings from innocent brand names in a language with a completely different representation system.

How a new EU eContent program can help you

0

Last week, the European Parliament voted in favor of the E-Contentplus program, set to support the development of multilingual content for innovative, on-line services across the EU. The budget will be about €149m for the 2005 to 2008 period. Program focus will be on improving the accessibility and usability of geographical information, cultural content and educational material in a large multilingual/multicultural marketplace. The underlying agenda is to boost broadband take up in Europe by making content better networked, more easily localizable and hence more attractive. Anyone interested in joining in had better bookmark this site and track developments.

The current eContent program (with a budget of only € 100 million) has funded market-seeding projects for access to public sector information, innovating production in multilingual/cultural environments, and facilitating the digital content marketplace (rights management). You can find out about them here. In the field of multilingual production, they cover such topics as subtitling systems (eTitle ) and ontology building for sharing legal information (LOIS).

These projects act as test beds for technology solutions, with a group of organizations working together on a real-world response that will eventually convert into a commercial service. In other words, the European Commission provides you with development money to grab a potential business opportunity.

Since this is public money, there should be more pooling of knowledge derived from these projects. For example, most of the translation and localization problems are solved in sui generis ways. There might be more interest in such programs if terminology, translation memory and other resources developed in these projects had to meet reasonable exchange standards, for example, so that some of the assets lived beyond the project in a collective way.

It is also nearly impossible to get any serious (i.e. drilling down below the usually thin website content) information about

(a) how the project went (intangibles such as what was learnt and shareable, what proved particularly problematic, what interesting ideas the whole effort spun off) and

(b) what happens to the tangible results.

When projects (such as eTitle) are packing in translation automation technology, speech recognition technology and the like, it would be good to seriously raise the learning level for the whole community by auditing these projects and providing feedback to the sector stakeholders rather than just to the Commission.

One solution would be to demand that a project Wikipedia be created, summarizing technical knowledge about the project in a semi-canonical form. Another is to develop a (possibly collective) project blog to make the work more responsive to their respective latent communities. There is naturally no need to reveal everything in the kitchen cupboard. But avoiding duplication, pushing the technology envelope and leveraging the existing infrastructure strike me as being basic desiderata for such projects.

Localizing propaganda in World War I

0

How much “official” translation was there and into how many languages in government departments and other power centers in the world before the League of Nations, the UN, SW and LW radio (the one-way global speech communication platform for the pre-digital generation) and finally the various modern institutions of global governance, made multilinguality a household word, as it were? Obviously it would have depended on your colonial spread.

During the 1800 to 1950 period, presumably France in its role as provider of Gallic clarté did little official translation into African languages (or Arabic), or any relevant Indian languages in its Chandernagore and Pondicherry outposts. One doubts that Germany did much in its short-lived Namibia, Cameroon and Togo colonies, nor the Brits in their African land grab. However, the administrative complexity of British Empire in a Brahmin-educated India raises more interesting questions about multilingual communication, and it would be good to have a thorough history of the linguistic relations between the colonizer (many of whose faithful Indian Civil Service henchman learnt local languages and legal systems to a high degree of proficiency, just as they had learnt the classics at school) and colonized (who have ended up learning various forms of local English in addition to their multiple languages).

Going further back to the Renaissance, diplomacy made use of Latin, French or other lingua franca, and nations usually tended as in the Spanish case in Latin America to extend the sway of their own language to the colonized (read Ivan Illich’s provocative Vernacular Values on all this). One imagines that the Vatican (or perhaps just the proselytizing Jesuits and other religious orders) did quite a bit of translating of religious texts and some official documents, even though Latin must have been a handy administrative esperanto for centuries Whether the Romans ever bothered to hire a 3rd century Visigoth, 2nd century Carthaginian or 0 CE Palestinian (presumably after Pontius Pilate had washed his hands, he wrote up his report in Latin) to localize Rome’s decrees would require research beyond the call of this posting.

What sparked this reflection was an item in the January 15th Times Literary Supplement (need to subscribe) by A.D. Harvey on the role of English writers such as John Buchan and Rudyard Kipling in the British ad hoc propaganda effort in the First World War once hostilities had broken out. I was struck by the amount of rapid translating going on to beam anti-Germany messages out to fence-sitters and neutrals. By June 1915, Wellington House (the propaganda admin center) had distributed “2.5 million books and pamphlets in 17 languages, as well as editions of the Bryce Report on German atrocities in 30 languages.” I calculate this report (written after the German invasion of Belgium in the autumn of 1914) to be about 30,000 words long, so it was only a three week job for a single translator, plus another three weeks, say, for typesetting and printing etc. But the chaps in charge – presumably unused to this sort of a multilingual workflow – were obviously capable of marshalling and managing translation resources (in Britain?) in war time.

Compared with the supposed language spread of old empire (and possibly the Bible translation agenda), 30 looks like a lot of languages for your average European administration to handle. Which were they? They would have at least included standard Western European, Scandinavian and Slavic languages (as well as Asian?) to reach that number. According to the same article, John (39 Steps) Buchan went to France and wrote a patriotic The Battle of the Somme: First Phase which was published in November 1916 and “quickly translated” into Danish, Dutch, Spanish and Swedish.

Any leads on “official” translation practice of this sort before say 1939 would be most welcome.

New organization to support translation automation users

0

The indefatigable champion of out-of-the-box thinking in the translation industry, Jaap van der Meer recently launched a Translation Automation Users Society (TAUS) to provide a professional platform for large-scale users of translation systems (from end-to end machine translation to workflows). Now that the web fields several dozen ‘consumer’ automatic translations services of varying quality and performance, and localization suppliers such as SDL are offering automated solutions for certain clients, it is clearly the right moment for user stakeholders to club together and share experience and best practices independently of any given supplier. 

Mood gear shift

0

I remember a language tech product that came out in France in the late 1980s designed to evaluate ‘attitude’ in documents. It used a dictionary to look up word connotations, and by combining connotations into attitude clusters, tried to identify the emotional stance – in fact the negative/ positive quotient – beating at the heart of document collections. It then used color coding of text to point out those chunks of text that exhibited the symptoms. I’m exaggerating a bit, frankly: the product actually started off as an automatic summarizer, but the inventor realized that the attitude of a document, its affective position on entities contained in it, were part of the ‘meaning’ and needed to be included in a good summary. Never heard of it again.

Today there are various technology initiatives underway designed to tap and exploit the emotional content of interactions, and possibly even of texts. At the most basic level, this is embodied in the sort of text mining found in the Reuter’s

Factiva project on ‘corporate reputation’. IBM’s WebFountain has now been dropped as the key technology, but the idea is to identify critical or laudatory remarks in say product reviews, or in analysts’ reports of M&As and so on to track reputations. You would need simple grammars of disdain and praise per language, and some sort of method of weighing up competing shades of attitude in a given search. Wanna know how X’s previous appointments went down in the trade press? Just mine the blogs and the opinion columns. Not perhaps affective computing, but the technology could have a field day with some of those spluttering me-too blogs or star-spangled SMS messages.

Human speech is obviously a much more sensitive indicator of personal emotional states, so it’s nice to see speech recognition technology getting another break here. This report shows how a Scottish firm Affective Media is using speech to identify car drivers’ feelings of road rage or drowsiness so that in-car solutions can be used to reduce them.

Affective Media chief executive Christian Jones said prototypes were being fitted to trial vehicles and claimed the system could be a life-saver. “Studies show unhappy or angry drivers are more prone to accidents than drivers who are relaxed,” he said. “Our technology will work with any voice recognition software. In the future, more cars will have voice-activated controls. This technology will sample the voice to tell if a person is angry or frustrated and will then act accordingly.

Alun Parry, spokesman for Toyota, said the company planned to test emotion-detecting technology in its experimental “Pod” cars. “We want a car to respond to the emotion of the driver and, as well as the voice technology, the Pod will monitor the driver’s pulse and could act to slow the car if it senses that the driver is being erratic or going too fast,” he said.

Inside the brackets

0

People delivering products and services in the area of translation, multilingual content management and the like need to be able to appreciate the real importance of the debate raging about the semantic web, which increasingly looks like a another nerds’ battleground. One problem is: how do we build the metadata and/or taxonomies – the bracketed tags – needed to help documents and services bind automatically and seamlessly into useful resources. Top down, through small technical committees who set standards or bottom-up through the vast distributed contributions of millions of users aided by simple aggregation techniques?

The solution chosen – perhaps the fudge of solutions – matters because language issues emerge at every step of the way. Will multilingual taxonomies need to be built, or will they emerge mushroom like from mega-scale group practices, prompting in turn the offer of further taxonomy aggregation services as a niche market? Should public money be thrown at large scale semantic web R&D projects that slide into obsolescence before they are completed, or should it be channeled into supporting focused language technologies for minority communities that tend to be crushed by the invisible hand of the market?

This comment is from Peter Norvig, Director of Search Quality at Google speaking at SDForum’s Semantic Technologies Seminar. He highlights the granularity of the language problems confronting web searching – it’s the data and whtehr it’s been spell-checked, not the metadata, stupid. Interestingly, he is mainly a bottom-up guy when it comes to solving the who does what issue of the semantic web:

Semantic technologies are good for essentially breaking up information into chunks. But essentially you get down to the part that’s in between the angle brackets. And one of our founders, Sergey Brin, was quoted as saying, “Putting angle brackets around things is not a technology by itself.” The problem is what goes into the angle brackets. You can say, “Well, my database has a person name field, and your database has a first name field and a last name field, and we’ll have a concatenation between them to match them up.” But it doesn’t always work that smoothly.

Here’s an example of a couple days’ worth of queries at Google for which we’ve spelling-corrected all to one canonical form. It’s one of our more popular queries, and there were something like 4,000 different spelling variations over the course of a week. Somebody’s got to do that kind of canonicalization. So the problem of understanding content hasn’t gone away; it’s just been forced down to smaller pieces between angle brackets. So there’s a problem of spelling correction; there’s a problem of transliteration from another alphabet such as Arabic into a Roman alphabet; there’s a problem of abbreviations, HP versus Hewlett Packard versus Hewlett-Packard, and so on. And there’s a problem with identical names: Michael Jordan the basketball player, the CEO, and the Berkeley professor.

Speechless in Europe

0

Is Asia edging ahead in the translation technology innovation stakes? News from Japan and Korea suggests that automatic speech translation (i.e. not written) will go live quite soon. And India is experimenting with language interpretation by radio. In Europe, 2005 kickoff news sounds distinctly underwhelming.

In the wake of the announcement that Korean mobile maker Pantech&Curitel’s latest handset features text to speech technology (whose, I wonder):

The TTS capabilities of the handset will convert incoming text messages to voice messages, and read them aloud as they arrive in loudspeaker mode. Also, users can control and have the tiny clamshell read the contents of its menu, phone book and unanswered calls.

and extends the fad for cameras with Optical Character Reader software to manage the local biz card ritual:

which can recognize text from JPEG files. This functionality is specifically aimed at those handling large numbers of business cards, which rather than requiring storage can simply be photographed with the built-in camera and parsed into the phone book of the device.

comes the suggestion in the Yomiuri Shimbun that Japan’s government-led “multilingual audio-text translation system” is likely to start testing soon and be completed during 2005. The plan is to offer a translation service probably for business folk to start with between Japanese, Chinese, Korean, and English over mobile devices such as phones and PDAs, using a database of ‘500,000 phrases or 5 million words’ (presumably meaning a million or so in each language). I can’t find any information on who is actually doing the development but suspect that ATR will be playing a key role for the ‘audio’ segment.

Meanwhile in Korea,

VoiceBuzz reports that two telecoms firms, SK Telecom and KTF, are to introduce in March two separate services that can translate Chinese, Japanese and English. Will this actually be based on the same underlying technology that will drive the Japanese government project? Although the German Verbmobil program looked well placed to deliver an operational speech trans system when it began in 1996, Japanese, Korean, US and EU speech translation teams have been working loosely together and sharing ideas in the C-STAR consortium. The ongoing European Union effort called TC-STAR for Technology and Corpora for Speech to Speech Translation started work more recently on new ways of building a talking translation machine. So while in Europe it’s been research, research and more research, it looks as if Japan and Korea – and via them China – will be the first to actually market an operational system.

In a democratic and massively multilingual country like India, access to linguistic services in public debate forums for people who can’t always read is now being provided by an innovative use of FM radio. According to this report, audiences at conferences can listen to simultaneous translations of the speeches by an interpreter, transmitted over a low band frequency from a control room in the conference center. This solution was invented by Hemant Babu and pioneered by members of Nomad, a grassroots organization that is developing tools for digital translation and archiving for non-commercial translation requirements.

The FM based translation system was first used during the World Social Forum 2004 in Mumbai, reputedly bringing down the cost of translations from US$223,000 to US$13,000. It was used again at Social Forum for South America at Quito, Ecuador in July 2004, and will be implemented next week at the World Social Forum at Porto Alegre, Brazil.

By comparison, gobsmacked you probably won’t be at the news from equally multilingual Europe that a project called TransType 2 , “an innovative computer-aided system that allows rapid and efficient high quality translations” is well…nearing completion. Using a mix of translation and machine translation technologies,

The TransType2 prototype is currently designed to assist translations from English to French, Spanish and German and vice versa, although additional European languages can be incorporated relatively simply. “To add Chinese or Arabic, for example, would take more research but it is possible,” José Esteban (project manager) says.

Translating the Dewey Decimal System

0

Peter van Dijck has an interview with the editor in chief of the Dewey Decimal System about translating the DD taxonomy. It’s worth noting that of the 35 translated versions of the DD translations, many strategic languages (Arabic, Chinese, Spanish, French, German etc) are either in progress or are very recent. This experience would provide an interesting case study in translating and managing multilingual taxonomies for some future semantic web, even though DD is a method for classifying books, not concepts.

Japan Prize to Nagao

0

Tip of the hat to Makoto Nagao for winning in the Information and Media Technology category of the 2005 Japan Prize for his “pioneering contributions to natural language processing and intelligent image processing.” Which is code for the fact he was largely reponsible for training a generation of young engineers in the late 1980s who developed all those machine translation systems in large Japanese tech firms.