EU factoids and figuresques

0

Anyone wanting a facts ‘n figures snapshot of where translation is at in the European Commission can find a start of the year summary here .

What you will not find is the back story to the 2003-2004 EU call for tenders for translation technology to possibly improve on some of the results given in this report. Apparently, the status quo was maintained and new technology suppliers were not taken on board. It’s a pity for the industry as a whole that it is so hard to find out what really drives decisions.

Patent news

0

Just before Christmas, Microsoft Research discreetly filed a patent for a data-driven type translation system: 

An adaptive machine translation service for improving the performance of a user’s automatic machine translation system is disclosed. A user submits a source document to an automatic translation system. The source document and at least a portion of an automatically generated translation are then transmitted to a reliable modification source (i.e., a human translator) for review and correction. Training material is generated automatically based on modifications made by the reliable source. The training material is sent back to the user together with the corrected translation. The user’s automatic translation system is adapted based on the training material, thereby enabling the translation system to become customized through the normal workflow of acquiring corrected translations from a reliable source.

But don’t give up on that project you had to beat the rest of them to market. Just this week, IBM announced that it would freeing up 500 patents for use by the Open Source Software movement. Included in the list, I found these covering aspects of automatic translation:

US5644775 Method and system for facilitating language translation using string-formatting libraries

US5251130 Method and apparatus for facilitating contextual language translation within an interactive software application

US5640575 Method and apparatus of translation based on patterns

US5267156 Method for constructing a knowledge base, knowledge base system, machine translation method and system therefor

US6236958 Method and system for extracting pairs of multilingual terminology from an aligned multilingual text

Others include

US5640487 Building scalable n-gram language models using maximum likelihood

maximum entropy n-gram models

US5636291 Continuous parameter hidden Markov model approach to automatic

handwriting recognition

US5220621 Character recognition system using the generalized hough transformation and method

US6249605 Key character extraction and lexicon reduction for cursive text recognition

US6182115 Method and system for interactive sharing of text in a networked environment

US5678052 Methods and system for converting a text-based grammar to a compressed syntax diagram

US6311177 Accessing databases when viewing text on the web

US6216102 Natural language determination using partial words

Can graffiti and public scribbling go digital?

0

In the Christmas edition of The Economist , there’s an interesting article on graffiti which concludes that the heyday of graffiti is over (partly replaced in the city by the pervasive anti-language of tagging) since the Internet offers a global outlet for ex-wall writers with a yen for anonymous antisocial messaging.

This is true in the sense that there are websites that offer megaphones to ranters of all stripes. But genuine e-graffiti would actually consist of messages surreptitiously splashed on someone else’s valuable real estate, not as message files to a chat room lodged on a receptive website. In other words, graffiti would have to be messages literally scribbled onto much-visited home pages (like popups?), not simply virus-carrying messages, email scams or hacker raids on databases. As far as I know, discursive popup graffiti have not been much experimented with. Presumably the security apparatus gradually being bolted into place on the Internet will soon make such irreverent e-scribbling technically impossible. You might argue that some blogs play a graffiti-esque role in the web’s community, but most blogs are still far too wordily hypertextual to rank as taut mural haikus.

The real digital problem with graffiti, though, is that part of their demotic charm comes from the iconicity of their interface – their mix of phrase and symbol (think Kilroy in Anglo-American culture) crudely inscribed by hand on public surfaces. Take away this physical “I was here” interface, and graffiti lose their punch. Until we get proper scribbleware (beyond the tyranny of the file, as Ted Nelson once put it) which allows us to write and draw at will on any sort of e-surface and have our scribbles saved and searched without having to name them, graffiti will have to stay on the world’s bricks and mortar.

What I’ve called here iconicity brings us to a further handicap for people who want to communicate by drawing. Without laboriously photographing your cartoons or sketches first and then saving them as files and then inserting the file in a blog or a website, it is very hard to share a ‘draw-pic’ (as opposed to a ‘photo-pic’) with others. Think how few bloggers or chat room messages, for example, make use of drawings to convey feelings or experiences. It’s just too complicated. For all the graphic ingenuity of web design, and the plethora of finicky graphics technologies devoted to geewizz image-making, we still don’t have powerful yet intuitive web tools for drawing a quick pic and scribbling a few words around it to make a point. So much for new media.

What about smilicons, you might say, as a replacement for personal drawing tools? I find them too standardized to express anything beyond a sort of metalingual smirk. So dear Santa, what I want for Christmas is a digital pencil box, not that digital camera.

Translation in the blogosphere

0

How can translation automation be tapped for the online world of blogs to ensure that information and ideas circulate freely in dark times? Tim Oren has been pushing the debate on this issue of ‘cultural bridges’, which brings together some of the threads in my own blog. The dilemma is clear enough and not new: are people prepared to work happily with machine translation output since it is better than nothing, or will the quality of that output and the attendant noise damage the actual conversation they are trying to have?

As far as I understand it, the background to this debate lies in the emerging power of blogging as social software and a personal publishing platform. There is something of that early web excitement about information trying to be free, of cyberspace being exterritoriality incarnate. See reports on the

Global Voices event and here for the broader concept of blogging as the new communication back-channel.

Now that millions of individuals around the world are being empowered, if that’s the right expression, by blog software to file reports from sites that the news media cannot always reach, or expose and share their personal lives in interesting new ways, or engage in intense discussion about issues of moment, far from the exclusive institutions of the nation state, the language barrier suddenly looms up like a wall of ice from the sea of words. How can we free the meanings encoded in those alien tongues?

There are of course dozens of free automatic translation services (that usually operate on very short texts) on the web, as well as a growing shelf-full of low price (but higher volume) solutions. But Oren seems to think that the blog community should have its own (open-source) translation solution, that avoids the awkward interfacing and quality constraints of existing web automatic translation solutions. His proof of concept, interestingly comes from the perhaps little known experiment on CompuServe in the mid-90s where users in the World Community Forum would converse both through and around the act of automatic translation, suggesting that where there is a will to communicate there’s usually an way round the problems.

In other words, is there an opportunity here to grow a grass-roots translation automation project? Precedents would include the infamous EU’s Eurotra project, which failed in its attempt to implement a total many-to-many translation infrastructure but nevertheless seeded the EU with a useful set of translation engineering centers for some 11 languages. Eurotra was largely predicated on a pre-web highly bureaucratic conception of project management. Perhaps a new automatic translation project for the blogosphere could be set up and managed as a Linux-type bazaar initiative, rather than a cathedral-like Eurotra. Especially since time is of the essence.

Closer to the spirit of freedom is the ongoing UNL project, which still seems to be known only to the cognoscenti, yet if successful, is set to become a genuine linguistic infrastructure for a networked world. UNL is the brainchild of Japanese experts and is being developed by a multitude of local language computational linguistics labs around the world. The UNL system is basically rule driven (statements are encoded into a UNL lingua franca semantics for later decoding into the local language), rather than data-driven (i.e. translations which are derived from statistical facts about existing databases of bilingual texts), and the language spread is heavily Asian-language oriented. UNL patented its technology earlier this year, so perhaps it would not suit the blogosphere’s mindset anyway. But it might be worth checking out whether the UNL protocols could slip neatly into the new blog space.

One interesting aspect of the UNL system is that the writing process itself can be adapted to enhance the translation process (known as dialog-based machine translation). This has the advantage of using human ingenuity to tweak text into something the universal semantic interlingua understands. Your blog would, as it were, be automatically translated into and stored as a hidden esperanto as you write, awaiting the call to be decoded into any available language. However, the work of encoding the spontaneous, fast-changing idiom of blogs and emails might defeat the very aims of the exercise.

Other projects that currently deal with parts of a potential translation infrastructure include the multilingual Wikipedia, whose article set could provide a multilingual resource for some sort of translation engine to work on, and the

Free Dictionaries project . This embryonic effort aims to produce a huge range of dictionaries through grass-roots lexicography, again driven by the availability of willing helpers working in a community.

Ideally, the development of such piecemeal dictionaries over many languages might be a useful asset to an online translation system, provide the databases are easily accessible as an online resource for different applications. Coverage will always be a problem (languages keep growing), as is lexical customization (a word is used in different contexts with different meanings – hence the essential need to disambiguate). But the plan to span very many languages might forge a handy learning environment in best practices for future generations of dictionary makers.

The standard criticism of this sort of amateur language work is that the data will lack proper quality assurance, that the information may not be correct, and that more fundamentally, open sourcing is not much good at bringing very complex projects to fruition. On the first counts, I imagine the best response is that the very openness of the approach means that in the end it acts as a self-correcting mechanism, and that enough eyes looking at linguistic data will eventually be able to iron out most errors and absurdities. As for organization, time will tell. If collaborative tools can be developed that help solve some of the organizational problems of very large horizontal projects, then some of these handicaps will disappear. The trick will be to persuade enough bilinguals to kick-start a project that is in a real sense the very paradigm of open sharing: after all language itself is the site of our deepest identity, as well as the prime vehicle through which we access the identities of others.

Multilingual scribes

0

There’s a meme going round about ‘amazing scribes’ at the recent ICANN meeting who transcribalated (don’t ask) spoken text onto screens. What attendees appeared to see were ‘scribes’ steno-translating at the speed of speech so that everyone could read a speaker’s translated content in real time on a large screen. Margaret Marks has rightly blown the whistle on this amazing feat: the steno-typists were in fact taking down text over the earphones from simultaneous interpreters.

Interestingly, the only languages used officially at ICANN were French and English, and the French translation facility was provided by the governmental Agence de la Francophonie which has a vested interest in a multilingual future for domain names and all that. But why weren’t more languages on offer at a time when ICANN is getting criticism about its ambitious and potentially expensive new Strategic Plan, and the fact that its mandate is to report to the U.S. Department of Commerce? In a world that is coming to doubt the exclusiveness of Lex Americana in these matters, showcasing a bit more of the amazing scribery of real time interpretation/ translation would perhaps help soothe the passions of multipolarity mavens.

As for the role of stenotypy in the communication process, those ICANN delegates who commented on the scribes appear not to have attended parliaments or assemblies where the record is always generated by steno-typists, court hearings of all sorts (the Nuremberg Trials were a pioneer in this, combining cabled language interpretation, rapid steno transcriptions and overnight text translation in many cases; but no big screen displays), and other venues where steno-typists (using rapid chorded keyboards) catch words on the fly and drum them into understandable text. Closed captioning on TV (where steno typists key in the speech stream) is another application area with a future, at least in the U.S. where there are legal measures for providing multi-channel language streams for the sensorially challenged.

Naturally everyone’s been imagining how to link up a translation rig to a steno-typist’s output to provide the sort of effect that wowed the ICANN watchers. This sort of platform would provide the kind of instantaneous translation that users of instant messaging systems or chat rooms have been dreaming of ever since the Internet came along. In fact CompuServe and others started introducing early instant online translation in the mid 1990s for constrained sorts of dialogs. As with everything else in translation automation so far, sometimes such dialogs work brilliantly but often they don’t.

Don’t forget, though, that many journalists and researchers using more than one language can take down real time conversations over a phone line in one language and type them into notes in another. Reader, I’ve been there. On a good day, you can input 1,500 words an hour this way, but you need to be a touch typist to avoid egregious errors. Applying this skill in public debate forums, you could probably keep it up for 45 mins a time (c.f. interpreters who sometimes do 30 min shifts); but most of all you would want to get paid twice – once as a steno-typist and once as a translator.

I’d be surprised if interpretation training course developers hadn’t already spotted a market potential in developing this ear-to-text trans-language skill in a wired world. Text tool makers could probably provide a few add-ons to accelerate auto-spelling correction were necessary, or helping in on-the-fly formatting. 

New language technology blog

0

Language Technology Business appears to be a brand new blog dedicated to tracking business news on language and speech technologies.  Nothing like a little competition to sharpen the quill.

Mind tapping

0

In France, there’s a court case underway to find out who, ten years ago during the Mitterrand years, was responsible for tapping the phones of public personalities who might have leaked the dreadful truth that the Prez had had an illegitimate daughter. Judging by recent reports, big daddy technology is now moving beyond human listeners in the quest for our secrets. According to the MIT Enterprise Technology Review :

Researchers from the University of Rochester and Palo Alto Research Center are aiming to allow computers to automatically assess peoples’ engagement in a conversation by analyzing the way they speak rather than what they say. The researchers’ system analyzes tone of voice and prosodic style, which includes changes in strength, pitch and rhythm.

Great stuff, but why?

It would be useful if a computer could sense ebbs and flows in conversation in order to automatically adjust remote communications systems. It would be useful, for instance, if a system automatically switched from a walkie-talkie-type push-to-talk system to a telephone-like full duplex audio connection when the participants become highly engaged in a conversation. The system could automatically adapt voice channels on-the-fly. It could also help a user who is engaged in conversation avoid distractions by deferring loud and new email announcements and changing instant messaging status to busy.

Sounds sooo user friendly, but natural users would be large government ears listening for the emotional rhythm of conversations between suspect citizens. After all theyet are already planning to spy in the textual medium. Smart Mobs quotes sources saying that the U.S. National Science Foundation and the CIA are to collaborate on research grants to develop “ways to monitor on-line chat rooms for terrorist activities.”

Professors Yener and Krishnamoorthy’s proposal, disclosed under the Freedom of Information Act, received $157,673 from the CIA and NSF. It says: “We propose a system to be deployed in the background of any chat room as a silent listener for eavesdropping…The proposed system could aid the intelligence community to discover hidden communities and communication patterns in chat rooms without human intervention.”

Yener and Krishnamoorthy wrote that their research would involve writing a program for “silently listening” to an Internet Relay Chat (IRC) channel and “logging all the messages.” One of the oldest and most popular methods for chatting online, IRC attracts hundreds of thousands of users every day. According to the proposal says their research will begin Jan. 1, 2005 but does not say which IRC servers will be monitored.

More on the FBI translator scandal

1

Useful udate from The Village Voice for anyone following the FBI’s (mis)management of its vital translation resource. It mainly centers on the perceived professional incompetence among people recruited through string pulling.

polyglot (the name) up for sale

0

ENLASO is selling (may have sold by now) its wonderfully evocative and brilliantly business-attractive polyglot.com domain rights, originally created in July 1995 and up for grabs in July 2005. Fools rush in…

Don’t you feel an end-of-an-epoch shiver about names like polyglot and babel? In an online world of powerful search engines, where multilingual domain naming is bound to go mainstream, and with multilinguality etching its way into the web’s DNA, those first generation real estate monikers sound oddly passé.

From a marketing perspective, polyglot will of course slip neatly into the language of advertising taxonomies. Which is presumably why it is going on sale. But in other ways, online habits are changing fast. As users, we go straight to the knowledge, not to website names – those quaint stopping points on the information highway, another sepia tint expression from the time when politicians were trying to leverage the old communication paradigm of driving cars into a glorious future.

What if search engines take over from websites as the key portals to knowledge and services? What if translation is increasingly carried out by pushing a translate button, not by visiting a translation supplier’s website? What, in a word, if online translation becomes a global grid function instead of a local choice? In this sort of regime, website names would lose their apparent value, and money from translation would be earned other ways.