Derestricting web corpus building

0

Language scientists and developers need corpora, but licensing them (no, not the linguists) can be costly. Either they are too expensive, or there are heavy restrictions on making versions or using them for commercial purposes. Or they add a heavy administrative overhead for gaining permission from all parties involved. In an effort to make it easier to build up corpora from existing web resources, Björn Lindström in Uppsala has come up with a ‘Creative Commons for Corpus Consruction’ which basically collects and parses web pages to check the metadata to see if they have an appropriate Creative Commons license. He’s found that the amount of material on the web licensed under Creative Commons licenses is “more than enough to build a large corpus”. The next step is doing something interesting with the corpus.

Wiki search engine

0

If you use Wikis, check out Wikiwax, a handy new index that uses the same sort of “predictive typing” tool as Google to suggest entries.

a-lettristic

1

Thanks to Language Hat for this piece of alphabetic fun by Simon Whitechapel. It consists of a set of rotating glyphs as stand-ins for Latin alphabet letters. Where possible, the glyphs distinguish minimal pairs (p vs b, etc) by the direction of rotation for the same shape. Vowels tend to be very simple shapes, consonants more elaborate. Drives you nuts.

The joke of course is that moving letters are illegible, even if you were to simply rotate the members of your normal alphabet. So they abolish the purpose for which they were created. Language aspiring to the condition of cinematics, as it were.

Pwnc: Cyfieithu peirianyddol

0

Among EU languages, Maltese is this year’s joker. Every time the European Commission announces a web site – such as this handy new Eurobook one-stop shop localized to 19 EU languages – they always add a rider that the Maltese version is coming soon. Yet with 400,000 speakers, Maltese is lucky to be an official language. Welsh has around 582,000, Irish about 800,000 speakers and Catalan… over 8 million. 

Official or not, Welsh is making doubly sure it doesn’t miss out on the fruits of modern technology – especially Cyfieithu peirianyddol, or machine translation. The Welsh Language Board, which oversees such policies, has recently launched an IT consultation, calling on interested parties to comment on a strategy document called Information Technology and the Welsh Language . If you need an overview of the MT options, read Harold Somer’s interesting report from last July on Machine Translation and Welsh: The Way Forward, which recommends a three pronged technical approach over two years to develop a useful automatic translation capability for the language.

A call for tenders to implement this recommendation has already been circulated to selected candidates, and the selection will be made after June 1 this year. Speech technologies are also part of the overall package. If this project takes off, it will be useful to keep careful track of costs, so that other languages with smaller speaker bases could benefit from good practice in this process. Maltese, for example.

Lack of French language technology?

2

First the call from Noel Jeanneney (see my post on this, and Mark Liberman for further well-informed and useful comments) for a European search engine to counter the Googlocracy in the domain of cultural content. Now a new complaint has come up about the lack of effective Euro knowledge management tools for competitive intelligence operations in France. The TNIS blog carries an interview (originally featured here) with Alain Juillet, a senior French civil servant in charge of business intelligence (BI), who reports to the General Secretary for National Defense. Sounds rather high level.

Asked what seems to be the problem with the French BI situation, M. Juillet replied:

Par exemple, nous manquons cruellement d’outils informatiques d’origine française ou européenne. Et notamment de solutions en matière d’extraction de données sémantiques ou vocales, d’outils de traduction automatique ou de moteurs de recherche spécifiques. Alors que maîtriser cette chaîne technologique est indispensable pour la sécurité et l’intégrité des transmissions dans le cadre du système d’information.

There is a worrying lack of French or European-sourced IT tools, especially for semantic or voice data extraction, automatic translation and specialized search engines. Yet it is essential to be able to control this technology chain for the security and integrity of communications within the (national?) information system.

In other words, you can’t trust anyone else’s knowledge mining and management (KM) tools when it comes to the highly competitive issues of business intelligence and, by association, national security. But should geopolitical suspicions translate so easily into technology options? Is it really possible for a non-Euro translation system to slip a subversive semantic bias into a given cross-lingual operation? The French seem once again to be saying yes: as with cold war nuclear weapons, KM is a strategic technology so it’s better to control it yourself.

Maybe this concern with the implicit “security” fear in U.S. tools is the real agenda behind the recent French Technolangue project that is currently trying to boost France’s overall language technology readiness. For the record, I wrote a somewhat snide article about all this a couple of years ago here. Yet according to a recently published Bureau van Dijk report commissioned by Technolangue (see my posting here), France has some 99 companies dedicated to language and speech technology, and a solid tradition of language processing excellence going back to the 1970s. Does M. Juillet know?

Obviously quantity does not mean quality. Yet ever since 1986, when Bernard Cassen (an academic, former Le Monde diplomatique writer and until recently president of ATTAC, the “international movement for democratic control of financial markets and their institutions”) wrote a whistle-blowing report on what would happen if Europe didn’t develop its ‘industries de la langue” for the emerging digital landscape, France can surely claim a pretty good track record in innovating in the language technology field.

France Télécom, which spun off good technology as Telisma, and LIMSI, one of the great European R&D labs for speech technology, have been world pioneers in speech recognition and text to speech technology; locally-listed Systran is a global translation automation brand name (and GETA in Grenoble is, among others, a major MT R&D center); and today there’s a slew of companies (Sinequa, Mondeca, TEMIS, Lingway) proffering advanced text mining and BI solutions.

So why is there still endemic concern about the lack of “semantic or voice data extraction, automatic translation and specialized search engines” when a quick phone call would theoretically net M. Juillet some of Europe’s best players? The key strategic wound that has never properly healed is the fact that France/ Europe today has no major computing company apart from…well, name one! Yes, SAP.

There is the concomitant fear that the Microsoft / IBM / Intel / Cisco road show is as much a threat as an enabler for Europe’s own IT knowledge agenda. Why? Because they run the desktops, enterprises and networks. Microsoft and IBM have also invested far more into research into KM type technologies (translation automation, and harvesting knowledge from unstructured text, be it written or voice) than most EU governments could even dream of. All of which means that the ultimate choice for M. Juillet and others will not between US and EU software solutions, but between IBM/Wintel and Open Source. But would France or the EU manage to control Open Source solutions any more effectively?

In memoriam Minitel

0

Jean Véronis has drawn attention to the fact that SMS spelling can be used to query the Yellow Pages (YP) in France.

Prsone n sembl lavoir remarqué, mè dps qque tems, les paj j0nes kompren le langaj texto: on peu cherché 1 6né, alle bwar 1 Kfé, etc.

He uses the example of ‘6né’ (which, being translated, as the Bible puts it, works out as follows: six-né = by pronunciation ‘ci-né’ = cinema = film theaters)

The input spelling system works largely on the basis of homophones for number words (1=un, 2 = deux, etc) and names of letters (k= ka, so a Kashmir restaurant will become kshmir…).

It also works for the English language version of the same YP website. I got the following:

‘2mbstones’ correctly returned “graves and cemeteries”

‘kidz’ delivered a variety of children-related addresses,

‘eats’ gave me restaurants and grocery stores, and

‘10sion’ more subtly got me doctors and scientific R&D agencies.

However, a cis-atlantic ’‘lorry’ correctly returned a page about ‘camions’, but a transatlantic ‘trucks’ only found an off-wicket ‘rolling stock’.

I have no idea how generalized such SMS-type search phenomena may be for Yellow Pages around the world, where English versions doubtless outnumber other languages whenever a local YP is globalized. But in the case of France, it is not simply the recent SMS trend but the singular history of Minitel that lies behind much of the language technology that enables the stenographic spelling queries that Véronis noted.

Minitel was a pre-web videotext system embodied in a small beige monitor with an ABC (rather than a more standard AZERTY ) keyboard delivered free of charge to millions of French phone subscribers in the late 1980s. It was initially designed as an interface to a massive database of French phone numbers in order to digitally replace the hugely expensive print version of the YP.

But almost immediately after it was rolled out, smart folk (remember William (Neuromancer) Gibson’s “the street finds uses for technology…”) realized that, telecommunicationally speaking, the Minitel protocol enabled network communications, not simply queries to a big database. Chat and dating services suddenly flourished like capotes anglaises around the trees in Bois de Boulogne on a Saturday night. And since users were paying France Telecom the price of a regular phone call (only dialup in those days), they inevitably invented a slimmed down French spelling system to make the time spent online more productive. Using abbreviations very similar to those described by Véronis today. frantic “Minitel rose” cruisers nearly two decades ago pretty much premiered much of the mobile keyboard idiom that has been reinvented today.

As far as I know, the French language engineering company ERLI (then run by Bernard Normier who now heads Lingway) was responsible for developing the language processing technology that deployed phonetics and semantics (thesauri etc) to deliver what at the time was almost certainly the world’s most advanced natural language interface for a consumer search product. Alas, by the mid 1990s, the global Internet tsunami had left France’s Minitel network stranded on an island of obsolete protocols. Yet this same technology managed to migrate successfully to the web and still drives the bilingual system we use today. 

Beyond tomAYto and toMARto

0

There’s an intriguing grass roots website that allows non-native English speakers to input their pronunciation of a test sentence which is then digitized. This feeds a growing database of speech accents of English that can then be used for various teaching, testing and other projects. Steven Weinberger (SW) of the George Mason University, Washington, has been masterminding it. Sound useful? Read on.

Why do dozens (hundreds?) of speakers with non-native English linguistic profiles read this test phrase: Please call Stella.  Ask her to bring these things with her from the store:  Six spoons of fresh snow peas, five thick slabs of blue cheese, and maybe a snack for her brother Bob.  We also need a small plastic snake and a big toy frog for the kids.  She can scoop these things into three red bags, and we will go meet her Wednesday at the train station.

SW: We constructed the paragraph so that it was short, had familiar words for non-natives, and had sounds that we wanted to test. The elicitation paragraph contains most of the consonants, vowels, and clusters of standard American English. To see the distribution of sounds, click here. We digitize the speech samples at 44.1KHz, 16-bit mono.

Do you keep statistics on how the archive is used and by whom?

SW: We don’t keep statistics, but we do know that more than 1 million visitors have seen the site. This is fabulous, given the arcane nature of the archive. This summer, we will be rolling out a completely new version of the archive, database driven and much more searchable. Then we will keep statistics on visitors. We get lots of mail from our visitors, who range from academics and engineers to stay-at-home moms. Everyone seems to like to listen to accents.

How about plans to use this database in news ways?

SW: We are letting the users of the archive exploit it as they wish, as long as they cite us as the source. We have a Creative Commons license, so people can use our stuff free of charge (so long as they do not sell it elsewhere).

We get lots of mail from speech engineers who use our recordings for speech recognition research, and from E(nglish) as a S(econd) L(anguage) teachers who are designing lesson plans from the samples. We even heard from a composer who was writing some saxophone music to the archive speech samples! We simply want to be kept informed of how people are using our stuff.

The Great Speech Translation Race

0

Article from Wired on U.S. progress in bringing the notorious Phraselator system used in Iraq back home for use by local cops. Here’s the current outlook for this technology.

The next generation of the devices will also feature pictures, allowing the user to ask, “Have you seen any of these people?” or “Have you seen these weapons?” The Phraselator is advertised on its website as an interrogation tool, but Sarich [Ace Sarich, vice president of VoxTec, a division of Marine Acoustics that developed the device] says it is inferior compared to human interrogators.

No kidding…

The article also refers to IBM’s MASTOR speech to speech translation project, which seems to have accelerated its development forecast:

In 2003, DARPA estimated that open-domain, multi-task and unconstrained dialog translation was still five to 10 years away. But the research group developing IBM’s MASTOR, or multilingual automatic speech-to-speech translator system, says its DARPA-funded bidirectional voice translator is a year or two from deployment.

This should be put in the hands of emergency workers in disaster areas, where more lives could perhaps be saved by having near-real time translation support.

South Africa speaking

0

There’s an article in the latest MIT Technology Review on South Africa’s effort to kick-start a multilingual technology industry from scratch. The idea is to address the combined problem of radical multilinguality (11 official languages), extensive illiteracy (which makes speech more useful than text as a medium in basic education, healthcare, the administration etc.), low penetration of IT infrastruture, and lack of interest among large IT suppliers (at 46 million people, too small a market for decent ROI).

The instrument being used is the Local Language Speech Technology Initiative, which emerged from the year 2000 Strategic Plan on human language technologies, built around partnerships with local and foreign R&D expertise. Read about the strategic vision for South African speech technology here.

Just this month the first downloads of text to speech software have become available in Kiswahili, isiZulu, Hindi (yes, an Indian language) and Ibibio (actually a Nigerian language). The idea is to stimulate developers to embed these text to speech applications in phone and other platforms, using a FOSS (Free or Open Source) approach modeled on that now used in India and Nepal. The key target: enable the local developer community to scale up and deliver the products in areas of the world where foreign big business would find investment too risky. This is a huge challenge, but South Africa’s efforts might well offer a useful best practice example of what can be achieved with a fully-fledged FOSS agenda.