European Patent Office MT system project

0

A recent post on the European Machine Translation mailing list from Jim Calvert of the UK Patent and Trademark office suggests that the European Patent Office (EPO) is embarking on a “very ambitious” project to use machine translation to speed up and lower the price of European patent processing. There’s no official information about this on the EPO site, but the word is obviously out. Two comments:

1. There was a lot of press coverage in 2003 of the linguistic problems at EPO (not an EU institution by the way) in their attempt to fulfill the trilingual translation requirements (ENG, FR, GERM) for certain patent documents filed with the organization. The problem was that the cost of translating patent documents made a Euro patent vastly more expensive than filing it elsewhere in the industrial world.

A Member of European Parliament said, “It is indeed true that the language problems are considerable. The average cost of a patent is about € 30,000, much higher than in the United States, and that is because 40% of those costs are taken up by language problems – the translation costs – and we are trying to get to grips with that problem.”

This from the life science industry shows how widely criticized EPO policy was: “The BioIndustry Association, the trade association for UK bioscience companies, has called for a simple and cost effective patent system for Europe to compete with those of the US and Japan. The existing European Patent system has its advantages but is expensive and complex, on average a patent for eight countries is estimated at € 50,000. This is largely due to translation requirements, as the claims of a European Patent must be translated into English, French and German at the grant stage. In most member states it must be translated into the language of that state to give the patent effect, with local registration fees added later.”

So a new project for an automated solution might be expected to lower the cost of translation in the longer term. It would be good to set some specific targets: say reduce translation costs by 50% (to 20% of the overall cost) in five years or so, so that people in Europol and other multilingual organizations with large translation throughput could gain from the experience.

2. Why is the invitation to join the project primarily targeted at universities? Europe has a splendid stable of MT companies with operational systems (Systran, ESTeam, SDL, Synthema, Linear B, Thamus, ProMT etc) which combine different approaches to the translation problem with a variety of language pairs. No one need be surprised about quality concerns in MT any longer – any solution would naturally need rigorous customizing and fine-tuning to address the highly technical domains in question. But it would be a pity if, instead of building a powerful, scalable infrastructure that could rapidly integrate these existing commercial suppliers, with their practical experience of rule bases, corpus tools and workflow, the EPO decided to start yet another very long term MT project from scratch.

Maybe they should first go and see Cross Language, Europe’s only fully-fledged translation automation integrator for a little advice!

yüerowz

0

European Central Bank President, the Frenchman Jean-Claude Trichet, has apparently told EU finance ministers that the word ‘euro’ was being spelt differently across Europe, thereby breaking a 1997 ruling that all official languages should spell it the same way. Well, not Greece, actually with its different alphabet – and in the future not Cyprus (EU ,but not eurozone) and Bulgaria (not even EU yet…).

Does it really matter that Slovenians write evro, Latvians eiro and Hungarians “euro with an accent”? We can already inscribe the neutral glyph € wherever the term euro has to go, presumably a sign invented after the name of the currency, and on a par with symbols such as &, @, the Arabic and Roman numerals. Oh, and until recently the sign for the singer known by the vocable Prince. All the rest, surely, is a matter for subsidiarity, that grand EU principle for trickling decisions down to country level. Different languages pronounce € differently, and now they tend to write it out in full differently. Some pluralize it, some capitalize the first letter etc, in line with their own orthographical or morphological norms. Indeed, the 1997 ruling, indeed, might well be the very first term in the EU dictionary that has NOT been localized in accordance with most other EU official language policy.

Cultural bias in taxonomies

1

Numerous bloggers, especially John Battelle, have picked up on David Weinberger’s dissection of bias in the apparently neutral Dewey Decimal system. As Weinberger says, any taxonomy will be skewed however universal it attempts to be, since its is intended for use in the real world, not in Plato’s world of pure ideas. Even programming a computer to compile a taxonomy from mining a large document base and allocate documents to apparently plausible taxons will just be one possible view of the conceptual scaffolding around that particular house of language. There could be many more, depending on what you are looking for.

This means that various forms of cultural bias will necessarily be built in to the taxonomies that will eventually underlie the semantic web. Multiculturalism itself is not so much the positive identification of multiple objective inventories of cultural ‘facts’. It is far more about the negotiations we make between different cultures, so that we find ourselves highlighting a specific feature (e.g. a marriage custom, an attitude to age, color, or embarrassment, etc) in another culture since it seems to be absent from ours. This mean that taxonomies that code and filter information will themselves need to be translated into other taxonomies to remove bias, rather as translators negotiate meanings between languages without relying on some angelic, universal language of concepts

Arabs and Arabic

2

Further to my blog on a new Arabic language translation center, .languagehat has blogged a Politics, Language and Cultures of the Arab World blog which seems to link to mainly English language sites.

Interestingly, the site’s headline term language is in the singular; yet surely there are more languages than Arabic (even with its various dialects) in the Arab world – Aramaic in Lebanese churches, Kurdish in Northern Irak, Coptic in Northern Egypt, Berber in Maghrebi countries etc. We don’t seem to use the expression Arabic-speaking countries – which would extend the geography to parts of sub-Sahara Africa – but do often tend to confuse Arabic the language with Arab (or even Islam) the culture.

I may be over-influenced here by French which has one word form arabe for the language and a person from the Arab world. But I notice a lot of slippage in the use of Arab/Arabic among English speakers too. Arabic is a language of the Arab world, albeit the dominant one; and the religious language of the Islamic world; but some Arabs are not Muslims and Arabic may be spoken outside of the ethnically Arab world.

Beit al-Hikma 2?

0

Good news from the UAE, where the Ajman University of Science and Technology has just set up a ‘state-of-the-art’ Linguistic Research, Studies & Translation Center .

One of the aims, it seems, is to address the needs for knowledge transfer by training a new generation of into-Arabic translators. Although there is plenty of translation education around the Arab world (certainly in the Maghreb and Egypt) there is little evidence that the results serve the vital cause of transferring today’s knowledge in the way that the Arabs in 9th century Baghdad translated and thereby kept alive ancient Greek philosophy and science, or in the 19th translated European knowledge at the behest of Mohammed Ali in Egypt.

The contemporary translation scene in Arabic-speaking countries recently received a blistering attack in the Arab Human Development Report 2003, (written by Arab scholars) citing various demeaning statistics on the number of books and other sources of knowledge that are Arabized each year.

While in the U.S. and presumably worldwide, there is enormous interest in mining information from Arabic texts and having it translated into English or other strategic languages, at least someone in the Arab world is trying to do the opposite and hopefully professionalize the local translation community for the more attractive task of localizing global knowledge for the future, not just local news stories for the very immediate present

Interpretation anywhere

2

One of cutest and presumably sincere soft-kill press releases that has come my way is an announcement for a brand new telephone interpreting service from Switzerland called Babel800 .

The idea is hardly original, but the Swiss Family Crélérot are cunningly cashing in on the ubiquity of the mobile phone to offer fixed rate fairly cheap interpretation services in some 100 countries via a virtual network of bilinguals who can work from home, in the train or down the local bowling alley for 35 € an hour.

Professional AIIC interpreters will probably rage against this new machine, but it looks as if Babel800 are pitching the service to a much broader spectrum of potential users than the usual interpretation clientele of cadres: truck drivers in Hungary, tourists in Timbuktu and police and firefighters working on air accident sites could all in theory get emergency linguistic help over the phone at any time of the day or night.

You’d have to be a subscriber, of course, which means setting up an account in advance on the web. But until someone puts a brain in Phraselator, or that pie-in-the-sky European Multilingual Companion is unveiled (see my Aug 2 blog and Robin Bonthrone’s comment), Babel800 type services will be a better bet than your PDA dictionary or phrasebook when talking foreign.

I say ‘services’ because the competition is bound to grow in this area, now that mobile phones are so abundant and can take photos of menus and signs that can be sent to an interpreter for immediate translation, etc. And Babel800 really ought not to live up to its name, especially on its English website. Yes, we have no ‘informations’…

Belgium 1 and 2

0

According to a news report from Brussels, Belgium’s new Foreign Minister, Karel de Gucht, suggested on Monday this week that Flanders (Flemish speakers) and Wallonia (French speakers) could function separately inside the European Union. This is not of course the first time that a Flemish speaker has suggested such a divorce for this endlessly bickering couple of communities. But now that there are more small-population countries in the EU, the arguments might appear more compelling.

It would certainly have the effect of radically reducing demand for Flemish-French translation, as well as normalizing the two ‘countries’ along the quasi one country/one language EU format. All countries are multilingual to some extent, of course, since they harbor minorities of all kinds, some more vocal than others. But explicitly multilingual countries would include Spain, Malta, Finland (Finnish and Swedish) and, plausibly, Ireland, which is trying more than ever to have Irish adopted as an official language alongside English. Minority language militants would obviously not agree with this interpretation of the linguistic space in Europe. To complicate the Belgium situation, for example, there is also a minority German-speaking community.

Yet if nothing else, an official split would at least rid each language group in Belgium of the publicly sustained illusion that they are expected to learn each others’ languages, and endorse out loud what each side thinks about the other sotto voce. The same is true for that other European multilingual country Switzerland, whose French speakers in Suisse Romande rarely learn Swiss German properly from their cousins in Zurich – and vice versa.

Paws up for the original Rex

0

A lot of people probably saw the announcement that speech tech developer Wizzard Software is entering its product Rex, the talking prescription bottle, in the competition for Best Embedded Solution at the 2004 Fourth Annual Speech Solutions Awards sponsored by Speech Technology Magazine.

Rex is a “self-contained, disposable prescription bottle that uses text-to-speech technology to read the medication information and instructions in a very natural computer generated voice.” It is designed to help blind and cognitively impaired users avoid making mistakes with their prescription drugs.

What the story did not say is why on earth Wizzard’s device was baptised Rex. Isn’t it an archetypal dog’s name, after all? Well, yes – and thereby hangs a tail. In 1922, the Elmwood Button Co. for some reason developed a one-off toy dog that recognized its own name. According to this source, “Rex was a celluloid dog with an iron base held in its kennel by an electromagnet against the force of a spring. Current energizing the magnet flowed through a metal bar which was arranged to form a bridge with 2 supporting members. This bridge was sensitive to 500 cps acoustic energy which vibrated it, interrupting the current and releasing the dog. The energy of 500 cps contained in the vowel of ‘rex’ was enough to trigger the device when the dog’s name was called.” If Wizzard did in fact know about good old Rex, why did they use it to name a device that offers speech production rather than speech recognition?

MS owns up to locale glitches

0

The localization community probably knows many of the venerable multicultural screw-ups made by various web designers when targeting communities and locales they simply don’t know about. But they are worth remembering for a new generation of practitioners. Jo Best of silicon.com reports that Microsoft’s top man in its geopolitical strategy team, Tom Edwards, revealed recently how “one of the biggest companies in the world managed to offend one of the biggest countries in the world with a software slip-up.” Apparently

When coloring in 800,000 pixels on a map of India, Microsoft colored eight of them a different shade of green to represent the disputed Kashmiri territory. The difference in greens meant Kashmir was shown as non-Indian, and the product was promptly banned in India. Microsoft was left to recall all 200,000 copies of the offending Windows 95 operating system software to try and heal the diplomatic wounds. “It cost millions,” Edwards said.

Another social blunder from Microsoft saw chanting of the Koran used as a soundtrack for a computer game and led to great offence to the Saudi Arabia government. The company later issued a new version of the game without the chanting, while keeping the previous editions in circulation because U.S. staff thought the slip wouldn’t be spotted, but the Saudi government banned the game and demanded an apology. Microsoft then withdrew the game.

The software giant managed to further offend the Saudis by creating another game in which Muslim warriors turned churches into mosques. That game was also withdrawn.

Microsoft has also managed to upset women and entire countries. A Spanish-language version of Windows XP, destined for Latin American markets, asked users to select their gender between “not specified,” “male” or “bitch,” because of an unfortunate error in translation.

Microsoft has also seen its unfortunate style of diplomacy have an effect in Korea, Kurdistan, Uruguay and to China–where a cartographical dispute saw Chinese employees hauled in front of the government.

Edwards said that staff members are now sent on geography courses to try to avoid such mishaps. “Some of our employees, however bright they may be, have only a hazy idea about the rest of the world,” he said.