Ludwig plays the lexicography game

0

The Guardian has a curious story about the recent sale for £75K of the proofs of a 42-page spelling ‘dictionary’ compiled by Wittgenstein in the 1920s when he was in non-philosophical mode. Listing bits of language to make a dictionary is of course a certain use of language, and in W’s later philosophy would presumably have been a type of ‘language game’, something you can do with words, a ‘form of life’ as he put it, with its own rules and institutions.

Dictionary making is certainly a virulent form of life on the web today. After the tedious grind, high cost and gross under-use of lexicographical products in the world of print, these days you can hardly start clicking without someone’s dictionary site popping into view. Now that almost any full word (not prepositions or articles etc) in any language offers an advertising entrée into some product or service, aggregating almost any dictionary resources looks like a money spinner. Still, few of them will ever reach the equivalent of 2000 bucks or so for one page of Wittgenstein’s Wörterbuch für Volksschule

Language: the 3D computer game

1

For all the sites for language games , constructed languages , artificial languages, Klingon, palindrome collectors, machine translation jokery, slang rangers and the like now active on the web, has anyone tried to apply decent digital design (DDD) to the task of showing how plain old natural language works?

By DDD, I mean deploying animation, graphics, and powerful multimedia interactivity. For the last couple of millennia, discourse about language (i.e. linguistics, translation theory and even language teaching) has made do with static 2D graphics, word lists, tree diagrams to show relationships between items, and various sets of icons to represent spoken pronunciation. We now have the computing power and associated design practices to represent codified knowledge about language so much more powerfully. Surely it’s time to pull language description as a discipline screaming and kicking into a visual culture of swarms, flows and 3D.

There are two obvious angles of approach for a radically dynamic visual representation of language. One would be to show how languages constantly change through time and space – the geography and history of all human tongues as interweaving kinetic patterns. This is the easy approach, since it is in a sense “external” to language as discourse. One simple attempt I have seen is this ‘animated’ site devoted to showing how various character sets, such as the Latin alphabet, evolved over time. But DDD I’m afraid it ain’t.

The other, far more complex line of attack would be to explore the internal structures and relationships of languages, again as visual displays of fields, waves and particles in constant evolution, not as lists of immutable rules. The problem here is that the analysis of language structure underlying such a presentation is theory-driven rather than an objective given. But you could surely generate a nice digital movie that includes all the different approaches to language analysis as they compete for explanatory supremacy.

The real issue about this DDD approach is that it would itself embody a ‘theory’ about language as flow, rather as our old bookish word culture of language led to a theory of structures that were somehow transformed over time. A dynamic, flowing view of language as endless coagulations of meanings emerging into symbolic form and then disappearing again, contrasts starkly with the mechanical plod of a history based on rules.

The linguist Andrew Wedel recently drew attention to this new linguistics agenda:

“I think there is a big shift from the explanation from a single level, advocated by Noam Chomsky, that one grammar algorithm is coded in our genes, to a more layered set of explanations where structure gradually emerges in layers, over time through many cycles of talking and learning,” he said.

“Languages are the ripples in the dunes and the grains of sand are our conversations, generations talking to each other and learning things and slowly creating these larger ripples in time.”

And Rob Freeman among other linguists is trying to develop computer models of language built out of rules that emerge from utterances, rather than utterances being pre-defined by rules. Now that we can stream our talk and text as bits and bytes that can instantly transform into any kind of multimedia pattern, we have a source of raw data whose own properties – not those imposed by a tradition of lists and rules – could inspire understanding. Or just entertain us.

Out of print in France

0

There was a sepia tinted article in Le Monde this weekend on the grim fortunes of the French Imprimerie nationale , founded by Richlieu in 1640 and thought to be the oldest traditional printing house in the world with functioning pre digital printing machinery for Western and ‘Oriental’ languages. Now forced to compete for jobs on an open market, maintaining the old machines as well as working to contemporary digital tech standards is becoming far too costly. Will they be able to preserve this national treasure?

Printing along with gunpowder and the compass, was classed by Francis Bacon in the 16th century as what we would call a ‘disruptive’ technology, a word that doesn’t even occur in Shakespeare. Add to that the steam engine and microprocessor, both of which have incidentally been used to drive printing practices in the last century or so. Today, though, printing has become just another display media, and its association with hot and cold type, and the nobility of craft work at the service of communication, will soon need an explanation in a dictionary.

Oddly enough, multilingual printing which we might spontaneously associate in Europe with the Dutch, has long been a minor French specialty. I live a few blocks from one of the shortest streets in Paris – about 5 or 6 meters long with only one doorway and no real number – but named ironically after the Abbé Migne who ran the largest printing house in Europe back in the 1850s during the Catholic renewal, and was responsible for some of the longest published series in Christendom – especially the Patrologia Graeca and Latina. Read about him in Howard R. Bloch’s very critical bio.

His Greek patristic series, for example, contains 168 volumes of dense double columned text, using Latin for the introductions and notes and the Latin translations of Greek texts, and is still found on library shelves, though the full-text database version is handier. The Latin series runs to 217 volumes. He developed a unique printing process whereby each text was proof-read 5 times, including a special ‘accent’ analysis for the non-Latin languages. He was a great believer in access tools and had 49 editors spending 500 man years making the indexes. By 1850s, his work represented 10% of France’s industrial output, with a staff of 300 printers, typesetters, editors and binders working in a Paris suburb. Someone had it in for Migne, though: The whole shebang eventually burnt down.

While we’re at it, Migne was succeeded a generation later by another unsung hero of multilingual technology. In 1910 Abbé René Graffin invented the first photostat machine. It consisted of a camera with a lens fitted with a prism, which would make copies on a role of light-sensitive bromide paper. In fact Graffin was director of the Revue de l’Orient Chrétien and founder of the Patrologia orientalis (in 25 volumes) and the Patrologia syriaca (featuring Syriac, Arabic, Coptic, Ethiopic, Armenian and Georgian). Hard to find out much about Graffin, but he certainly pioneered the use of photo-copying of documents in his efforts to publish the content of the Levantine manuscript tradition, and even designed the type faces for the printers.

Alternative translation networks

0

For anyone interested in how progressive groups such as the European Social Forum handle the political minefield of multilingual communication (all languages are equal, all participants are equal, etc.) without a real translation budget, read Babels and the Politics of Language at the Heart of the Social Forum here. It offers fascinating parallels to the kind of problems met in the hierarchical, commercial world, both at the level of professional ‘IP’ concerns and in terms of technology fixes – in this case the development of what we might call corpora for later terminology mining. And it suggests that translation management is by its very nature a problem site (rather than just as a bunch of plug-in, goodwill-driven services) that raises issues that go way beyond strings, codes and words. But I’m surprised that there is no mention of links to the open source movement, which might provide a potent technical resource for software responses to Social Forum needs.

The article refers in particular to the work of Babels, the organization of interpreters and translators, and NOMAD, an umbrella group that tries to put technology to work for this translation space:

Babels is developing innovatory new language tools through activities like the Lexicon Project. This is an on-going effort by volunteers from a wide range of countries and backgrounds (teachers, students, professionals, activists) to create a comprehensive glossary of words and phrases to help interpreters and translators best reflect different meanings according to different national, cultural and politico-historical contexts. It is consciously creating a process of ‘contamination’ in which the excellent language skills of the politically sympathetic trained interpreter/translator interact with the deeper political knowledge of the language fluent activist to constantly improve the communications medium within the Social Forums.

Lexicons are being formed in conjunction with the Situational Preparation Project, more commonly known as ‘Sitprep.’, which records WSF and ESF plenaries and seminars in a wide range of languages on to DVD to allow any volunteer – experienced or inexperienced – to more realistically prepare for simultaneous interpretation in the Social Forum. This issue links to the broader ‘memory’ implications of the NOMAD project to which Babels belongs. As Sophie Gosselin argues elsewhere in this newsletter [I couldn’t find it], one of NOMAD’s main achievements so far has been the creation of Targ, an open source software system which can replace expensive propriety audio equipment used for live simultaneous interpretation. In addition to the revolutionary cost implications, using computers to relay the voices of speakers and interpreters the Targ system enables all speeches and interpretations to be easily archived, creating a direct and accurate ‘memory’ of all the debates, themes, and controversies of each Forum. Taking it a step further, the audio could also be streamed live over the Internet. These possibilities would allow millions of people currently outside of the Forum to take part via the web.

Significantly for Babels, the creation of Memory will allow the quality of interpretation to be assessed and new online ‘distance practice’ materials for inexperienced volunteers to be created. Not everyone will welcome this latter development within Babels. Many interpreters are already reluctant to have their work scrutinised and not just because they are ‘volunteers’. While professionals are simply not used to such practices in their particular labour market, non-professionals are often worried about being judged badly and marginalised. But if Babels is genuine about its commitment to ‘equality’ and ‘quality’ of communication’ within the Social Forums, then these worries will hopefully disappear.

Report on the European Language Tech Industry

0

Under government auspices, a new report has been published by French information and language industry bodies on “Automatic Language Processing in the Information Industry”, researched by the Paris office of Bureau Van Dijk. Well, maybe not such a new report: the data comes from a 2002 survey and I believe was first presented at LangTech 2003.

Here are the key findings as presented in a Paris press conference today:

· The language and speech tech market divides into 80% text and 20% voice. [Not everyone would agree with this].

· The European LT market in 2002 was worth 510 million €, or 33% of the world market.

· UK, France, Germany and Italy accounted for 60% of this market, the UK having the biggest share at 18%, largely due to its flourishing speech tech market.

· In France, the key segments (from document management via translation automation to searching) were content management (search) type applications (25%) and document management (categorization etc) systems (20%). Translation accounted for only 10%.

· Most players estimate growth –to stand at 10-15% a year. Most demand is for document management at 23% and search applications at 21%.

· ROI come mostly from time saving rather than financial returns. Main costs incurred come from feasibility and needs analysis, followed by software development, then customization and maintenance. Multilingual applications in general double the ROI horizon.

· Main trends: shift from supply to demand, as new mobile and desktop platforms in a multilingual search and personalized work space require specific tools.

· The European LT market is growing at 6.9% annually, compared to 9.4% in the USA. The global LT market will be worth 2 billion € in 2005 and should rise to 3 billion in 2007.

I don’t think these findings will surprise anyone. Language technology appears to be but a minor segment of the software or IT industry, with lots of small players (over 300 in Europe alone) that have emerged as spin-offs and startups from the R&D base and are now ripe for M&As. Consolidation is already more striking in the U.S. than in Europe, probably due to such structural brakes as nation/language-based fragmentation in Europe, and an endemic culture of lower investment in risky ventures. This means that current players tend to have small customer footprints in a few selected industries where advanced business intelligence (rather than, say, translation) is a key concern. The scope for real growth and infrastructure investment is often minimal

But if an ambitious global player in a related but external segment – voice technologies, for example – decided to integrate language technologies to provide a better user experience for, say, mobile phone/PDA access and search with built-in multilinguality, then that player might start cherry picking any European players with good technologies for translation, or semantic searching. It happened once before in Flanders Language Valley, though aided by considerable public investment. For the time being though, it looks as if pervasive language technology will continue to be “decades away” as press articles keep telling us.

Lesser spotted MT specialists

0

Friday trivia. In one of the most extraordinary bursts of messaging I’ve ever seen on the MT list, an online forum for machine translation specialists, the eminences grises of the profession, who never usually contribute to ongoing MT threads, all reacted with unusual alacrity to a minor query from Harold Somers of Manchester University: grammatically speaking, shouldn’t you say “less used languages” in English instead of “lesser used languages”? While machine translation topics are usually restricted to PhD students asking where they can get a corpus of this or a system to do that, less-ness was suddenly more.  This must be a case of ovian mimeticism – once one elder statesperson puts finger to keyboard, all their peers feel obliged to join in. The grammar issue was not solved. And MT naturally never came into it.

Updating Kanji

0

Useful item from Burritt Sabin on a recent Japanese Language Subcommittee report on kanji, calling for an overhaul of the joyo kanji (the official 1,945 characters used in administrative life and education).

…personal computers, cell phones, and other IT devices may be encoded with as many as 6,355 characters, the total of the first and second levels of the Japanese Industrial Standards (JIS). Simple arithmetic shows the Japanese are using a lot of non-joyo kanji.

What’s more, the use of Japanese and Chinese readings and character styles not in the joyo kanji list is increasing. For example, the old forms of the kanji for “kuni” (country) and “sawa” (valley) have come to be widely used in personal names. As well, many kanji used in place and personal names are not in the joyo kanji list. The reason, according to a Cultural Affairs Agency spokesperson, is that “proper names were expected to be handwritten.” The joyo kanji list was intended as a guideline for printed characters. Even “saka” in Osaka and “oka” in Okayama are not joyo kanji.

See here for more on Japan’s language policy (in English).

Patent searcher

0

Further to my post on IBM releasing numerous language and speech tech patents to the open source (OS) domain, it’s worth mentioning that The PatentCafe’s Patent Search Engine is now offering search and delivery of PDF copies of all the 500 IBM patents for a nominal fee to the OS community.

Recent translation industry research

0

Common Sense Advisory (CSA) has published its Fourth-Quarter Global Business Confidence Survey on the translation services industry.

The surveys polled buyers and suppliers of language services and technology about their current business situation, plans, and expectations for the near future.

CSA has also just published SDL 2005: Challenges and Opportunities, the second in its series on assessing the (few) publicly traded language service companies.

After IDC and ABI Research and a few others had a go at mapping the language services / localization (software) space during the late 1990s and early 2000s, CSA is now establishing itself as the only research company that is providing sustained close tracking of the translation / localization industry in the context of global business challenges. I should declare an interest here, since I have worked with CSA, though not on these reports. Still, since you only work with companies because you like the way they do things…