New translation blog

0

Gabe Bokor has started a translation blog on the revamped Translation Journal site that should interest freelance translators.

It is oddly called ‘interactive’, since the declared idea is to start discussion rather than report on the world from the blogger’s viewpoint. Blogs can obviously generate discussion, but to my mind, the ‘comment’ function does not work nearly as efficiently as chat room type interaction where everyone is engaged with a topic. Time will tell.

U.S. report on FBI Language Services

0

Further to my tongue-in-cheek blog the other day concerning the news about the FBI’s 120,000 hours of unread phone tap content, the report from the Justice Dept on the Bureau’s performance in language management can be found here.

yakushite distributed MT service

0

OKI Electric has finally launched its distributed automatic translation system as a free service for Internet users. The heart of this system is the nearly 20 year old PENSEE MT (Japanese<==>English) system, rewritten in Java at the end of the 1990s. The big news is in fact about one crucial design feature of yakushite – it helps user ‘communities’ to enrich the system’s dictionaries. The idea of grass roots word wallahs uploading reams of likely lexemes into dictionary warehouses for free must have crossed many a resource-poor language tech developer’s mind since the Web unfurled. It would be interesting to know how many have gained anything really useful from actually doing it.

One classic example is the Italian language services vendor LOGOS’ online Dictionary which has been fed by individuals to the tune of seven and a half million (validated but not guaranteed) terms in an unequal mix of languages. It would be a good thing for someone to take a close look at LOGOS’ dictionary and reveal some of its internal ecology; there may even be marketable tools around to find out a) what the distributions of languages and term fields are, and b) whether resources like this can be ‘X-rayed’ quickly for potential users. Of course LOGOS owns the database containing the contents, so this won’t happen soon. But appropriate analytics would also allow LOGOS to add value to its resource.

In the case of yakushite, OKI has wagered on enabling subject matter ‘communities’ to build dictionary resources to improve output quality . Apparently it has developed some user-friendly term extraction tools that speed up the candidate term compilation process. The idea sounds great, but will people actually want to spend much time extracting, checking and uploading a subset of terms about their small terminological corner of the multilingual forest just to translate a travel website? One obvious place to try this would be in schools and universities, where subject matter is a constant focus, and where there is plenty of cheap labor to extract and prepare dictionary resources in exchange for the odd credit. But it is difficult to imagine many professionals from business sectors sharing their high specialized, patchy yet vital glossaries with OKI.

Bengal Bill

0

The Bengal Observer reports that Bill Clinton’s My Life is being published today in Bengali, the language of Bangladesh. Interestingly the report says:

The book has created sensation in the world recently. The book is the first ever Bengali translation in the country but also in this sub-continent.

The noted scholars and personalities responsible for the editing and translation of ‘My Life’ were Akimur Rahman, Dr. Nurul Haque, Shaheen Reza Noor, Mosharraf Hossain, Mohammad Shahidullah, Shamsur Rahman, Mahmood Menon Khan, Kamal Morshed Milton, Sarwar Hossain, Mostahid Hossain, Bisjwajit Basu and Abu Sayeed Khan.

I presume that ˜first Bengali translation” claim simply means that this Bengali translation is the first one made of Bill’s book in Bangladesh and in India Bengali being widely spoken in Calcutta, and other regions of India. There have surely been previous Bengali translations of other books, especially if Bangladesh can dedicate so many scholar translators to the task.

FBI yet to decipher secrets of translation

0

There are news items all over the map today on the U.S. Justice Department report that there are more than 120,000 hours of “potentially valuable terrorism-related recordings” that have not been translated by FBI linguists due to backlog, management and “computer storage” problems. Apparently a lot of the report’s criticisms of the translation effort remain classified.

“What good is taping thousands of hours of conversations of intelligence targets in foreign languages if we cannot translate promptly, securely, accurately and efficiently?” said Sen. Patrick Leahy of Vermont, the ranking Democrat on the judiciary committee.

“The Justice Department’s translation mess has become a chronic problem that has obvious implications for our national security,” Leahy said.

Let’s try and help the FBI with some creative accounting on those figures, and show that there are plausible translation solutions, even though they have failed to detect them. If we agree that speaking speed is a traditional 250 wpm, and the FBI has a backlog of 123,000 hours of the stuff, this makes around 1.8 billion words in a mix of Farsi, Arabic and other languages, that need transcribing and then translating (possibly selectively) into English.

Let’s say you wanted to do the job properly: first transcribe the whole of the spoken record into electronic text, and then edit and ready the master file for information retrieval and/or translation applications. You manage to put together a transcription team of 50 trained Arabic language stenographers (using chording keyboards for speed). They would each have to take down some 36 million words, and working 8 hrs a day they would be at it for over a year. Let’s say you pay around US$ 3.5 million for the transcription and subsequent editing.

If you can’t wait for the stenographers to get their act together, you can always try transcribing speech signals into electronic text by using a speech recognition system with language model capabilities for Farsi, various dialects of Arabic, and any other languages, etc. Currently, streaming this audio signal through the kind of cutting-edge 2004 products that were demonstrated and discussed at the recent SpeechTEK in New York, you might achieve a 20% error rate (this is being pretty generous, given the spotty quality of the input signal) in the transcription, making it useless (since incomplete) as an intelligence source. Cleaning up and editing the output would be time-consuming and costly.

But if you did go down the speech recognition path and came up with a reasonable automated transcript, technology currently on the market would be able to audio search the sound file for at least names, places and dates to see what’s in the transcripts. See HP’s speechbot or Streamsage among others. This would enable relevant parts of the transcript to be sent on for translation into English and hence to enter the FBI’s knowledge radar.

Now suppose you had an automated translation solution that would deliver a gisted version of the original Arabic / Farsi etc audio tapes – plenty of errors but lots of actionable content, as they say, between the errors. With a translation throughput of say 100,000 words an hour, you could theoretically translate the whole transcribed shebang into English in about 75 days of non-stop processing. And obviously in a fraction of that time if you need to just gist a few selected passages. If the right language pairs in available systems were operational, you could negotiate an MT price of say 5 cents a word, and your translation budget for this backlog would reach US$ 1.53 million. Let’s say the whole backlog cleanup costs around US$ 12 million.  Its a winner. According to the new reports, the FBI language services have a 2004 budget of US$ 70 million.

This is not the first time large scale phone tapping has had to be recorded and translated for intelligence work. David Stafford in Spies Beneath Berlin recounts how in Cold War Berlin in the early 1950s, the UK intelligence agency MI6 working with the CIA, built a tunnel called Stopwatch/Gold that ran for 1,924 meters under the Soviet Sector of Berlin in order to tap into sensitive Soviet phones. The predominantly Russian language calls captured in this way generated 25 tons of 2.5 hr tapes. Let’s say a tape contained around 37,500 spoken words, and the physical object weighed around 12 ounces; the phone taps added up to about 80,000 tapes, and a total audiostream of about 3.4 billion spoken words. At one stage, MI6 had to translate 20,000 of these tapes containing 368,000 communications or 75 million words using a staff of 300 transcriber/translators. We can only hope the war on terror ends, like the Cold War, with a whimper rather than a bang.

Can you speak German at the U.N.?

0

At the weekend, e-world thinker Esther Dyson criticized the United Nations efforts to work unilaterally (via a new Working Group on Internet Governance ) towards regulating the Internet as a global ‘facility’. In a nutshell, Esther believes the Internet should not be regulated by any single centralized authority such as the U.N., even though it could play a role. Users everywhere should have a say in how it is governed. Which means people speaking different languages meeting together to thrash out strategy. One argument against centralization came from Adam Peake from the Centre for Global Communications in Japan, who is quoted as saying:

“Everyone must be able to participate, if not in person then remotely or through submitted comments, and that should be in their local language,”(…) pointing out that languages such as Japanese and German were not ‘official’ U.N. languages.”

In other words, if a given international body takes over the control or ‘governance’ of a global facility or activity, its own linguistic regime will determine its accountability to citizens. But is this actually true?

At a trivial level – yes. The U.N. certainly excludes German and Japanese, along with some 5,900 other languages on the planet, as official working languages. As it happens, there is a grass roots movement to have Japan, Germany, India and Brazil elected as permanent members of the Security Council, which would entail official status for Japanese, German, and Brazilian Portuguese if not an indigenous Indian language. In fact, Joschka Fischer spoke at the U.N. in German the other day.

But international organizations (as do some corporations) always make a difference between official/working languages and actual face to face needs. If a Japanese speaker wishes to speak in some capacity at the U.N., I imagine they would be provided with two-way oral interpretation or translation facilities, as appropriate. The OECD in Paris has only English and French as its official (working) languages, but its translation department handles a vast range of languages beyond these two, depending on specific needs. A German monolingual speaker could in practice, I believe, communicate to committees and groups in these bodies, however awkward it might be.

Likewise, the E.U. has a small group of ‘legacy’ working languages for day to day meetings, to the consistent dismay of newer Greek, Portuguese and Danish delegates (and now another 10 speaker groups). If necessary, though, Brussels will roll out its whole linguistic works for high value individuals, such as Turkish prime minister Erdoğan who has been communicating his country’s desire to join the EU in Turkish. I also expect that any of those organizations mentioned would field signers and produce Braille transcripts had someone like Beethoven or Helen Keller turned up to complain. An organization’s linguistic policy probably does not determine its informal communicative repertoire, only its official face. But it might be worth checking on how far this is possible all around the globe.

What’s in a wor(l)d?

0

Blogged by John Battelle, the interview with Ramesh Jain in the latest Ubiquity reveals some interesting ideas about the ‘future of search’ meme, but it also contains an unintended warning to people who don’t re-read spoken (as opposed to emailed) interviews, or expect editorial support.

Here’s the passage, which is, of course grin, about language.

UBIQUITY: Would I be wrong in saying that the essence of what you’re doing is trying to get beyond language?

JAIN: That’s very correct. I think one of my favorite books, which I love and quote a lot, is a famous book by S. I. Hayakawa, “Languages (sic) in Thought and Action,” in which he makes the point that it’s amazing how some arbitrary noises and some scribbles on paper started to make meaning to us. It’s amazing that we all agree that some noise is going to be presenting some particular thing or some particular concept. And similarly it’s amazing that we agree that such-and-such particular scribbles are going to be representing this and that real thing. Another person who had great insights on these issues is Carl Popper. Popper started talking about word one, word two, and word three. Where word one is the real word, word two is what concepts and what mechanisms you learn and have in your head, and word three is the model you build using word two about word one. And those things become very exciting because what we don’t see — which is word two (what we have in our head) — is a lot more complex and sophisticated than our language allows. And that’s how we sometimes fail to find the words to represent what we want to say. Language allows us to represent some of the things that become a lot more explicit and a lot clearer. Language is a knowledge representation language.

I guess Jain was speaking live since the transcriber got Carl wrong (it’s Karl), but more importantly managed to hear Jain’s utterances about Popper’s World 1, 2 and 3 (philosophical concepts) as word 1, etc..

What anyone understood from “word three is the model you build using word two about word one” is anyone’s guess. Maybe it was because they were talking about language that Jain/Popper’s ‘world’ morphed (almost understandably if you didn’t catch the reference) into Ubiquity’s ‘word’. Maybe it was because of Jain’s pronunciation. Whatever the cause, the result is incomprehensible. And the moral must be, make sure you design and implement an editorial process if you don’t understand what your interviewee is talking about. This should include at least a ‘speaker’s cut’.

Just for the record, Karl Popper said:

“If we call the world of ‘things’ – of physical objects – the first world, and the world of subjective experience the second world, we may call the world of statements in themselves the third world (world 3)… I regard books and journals and letters as typically third world objects, especially if they develop and discuss a theory… I regard the third world as being essentially the product of the human mind. It is we who create third-world objects.”

In other words, World 3 is surely what we now call content, the stuff we increasingly use machinery to process As Popper suggested:

“human evolution proceeds, largely, by developing new organs outside our bodies or persons… instead of growing better memories and brains, we grow paper, pens, pencils, typewriters, dictaphones, the printing press, and libraries…the latest development (used mainly in the support of argumentative activities) is the growth of computers” (Objective Knowledge: an Evolutionary Approach, 1972)

Pronunciation (pronounced /prer’NUnsi’EIsh-n/)

1

Mark Liberman’s learned Language Log post on Italian pronunciation and how The New York Times journalists get it wrong inadvertently draws attention to a real enough communication problem: how can we use media to tell others how to pronounce words?

Liberman imagines a parallel universe in which NYT journalists would come to know the International Phonetic Alphabet, with its large repertoire of unique symbols for unique phones that cover the sound elements of any human language. Some hope. Closer to reality is the Merriam-Webster online dictionary(for English, natürlich) approach which offers a non-technical but still awkward pronunciation guide (e.g. ‘gId where the identifies the stress onset and the rest, of the symbols may have to be looked up in another click-intensive operation) for each word in the language. For a fee, M-W also offers an audio version, which presumably allows you to hear that the word conundrum is pronounced /ker’NUndr-m/ and not /‘KOn-n’DRUm/, which is how I used to pronounce it to myself as a readerly kid. But whether the audio service uses a male or female voice or allows age and geography variation in the speaker, I cannot say, being too mean to pay.

One obvious problem with existing dictionary multimedia aids like this is that they inevitably miss a vast range of real needs about how to pronounce proper names, nonce words, technical terminology and evolving symbols in our everyday encounters with written language. As an example, here’s a googled selection of how web writers try to help readers with pronunciation :

· “Balluchillish Pronounced “ball-a-hoollish”

· “Jonathan Laughlin pronounced lock-lin”

· “Logic may dictate the “g” in GIF (Graphic Interchange Format) is pronounced hard, like gift or gefilte fish, but it’s “Jiff” and I Don’t Want to Hear Another Word

· “C# (pronounced See Sharp ) is scheduled to be included with the next release of Microsoft’s Visual Studio .NET programming environment”

· “VLDCMCaR (pronounced vldcmcar) Very Large Database for Concatenative Music Composition and Recontextualization”

· “HAOLE presents (pronounced HOWL-ee)”

· or this on how to pronounce the @ sign in French (the people voted for arobase pronounced /’ARro’bahz/ whereas the French government terminology commissions says “arrobe” or /arr’Ob/.

Presumably text-to-speech technology (TTS) will come to the rescue in the long run, but it will be an expensive business to speechify all words in all languages. Instead of linking to a pronunciation site, you should be able to find an on-the-fly TTS service embedded, as it were, beneath the word on the screen, rather like the various clickionaries that pop up a definition. (indeed, I imagine that web users/citizens will have this sort of access to any amount of automated linguistic and encyclopedic knowledge in the coming years). For on-the-move pronunciation angst, we should be thinking about a mobile phone service you can ring up to get a pronunciation prompt for a hard-to-say name from your directory, just before a meeting, or for the name of a town from a list. But memory can betray us, and literates tend to write pronounciations down in their own way. Which brings us back to reading. When you are reading text – newspapers, maps, address books, name lists at a conference – there should still be some plausible method (language specific obviously – no universal values assigned to Latin alphabet symbols, please) for sketching an audio image into the reader’s brain using familiar symbols, not phonetic science.

The oddest thing about pronunciation aids in a connected world is that the written ‘pronunciations’ will all have to be localized: Hungarian names will have to be inscribed in some appropriate local form for Spanish, English et al. speakers, just as Spanish will have to be for Han, Wolof et al. speakers. This way madness lies. Maybe we ought to just barcode the whole language and as the reader is passed across those pesky words we don’t know, a neutral pronunciation will go straight into our earphones…

MT prizewinners

0

Tip of the hat to MT companies who have recently won or helped win prizes for their enterprise solutions. On Monday the Depart of Motor Vehicles in NY State announced that it won the 2004 American Association of Motor Vehicle Administrators Region I Customer Service Excellence Award for on-the-fly English to Spanish translation of its web site using SDLEnterprise Translation Server.

And Systran has just been awarded an Information Society Technology prize by the European Commission for its Systran Professional Premium translation solution. If it wins the next round, it might earn an extra € 200,000 as a Grand Prize winner, but will have to wait until early 2005 to find out.