Bridging the common language gap

0

Churchill quipped that the US and the UK were two countries divided by a common language. Despite a two-way stream of mass media products and mutual exchanges, this divide is still a constant operational issue for readers, writers and translators. It involves such simple to grasp yet pervasive components as lexical items, phraseology, spelling and punctuation. Many a European based translator into English ritually checks with a client to know whether they want US, UK or just plain “international English”, which is usually the translator’s natural idiom beefed up with ‘z’ spellings for words like organization or recognize. But nowhere does the growing population of online media readers help than in the realm of slang: witness this comment from a US reader of a UK IT news publication that likes to sling the latest Brit slang around for maximum effect.

Rapid and ubiquitous web publication has naturally exacerbated the problems that come from thrusting US and UK noses into each others’ linguistic troughs. And sure enough, there are web solutions now offering a growing range of look-up solutions to this common language problem for US-UK slang. Transatlantic partners are not alone. Even Oztralian familiar from sitcoms and nomadic media-smart Aussies poses a problem to many viewers and readers. Most of these online services look decidedly laddish (a Briticism you can check here) but the tsunami of blogs, where ‘natural’ idiom is a defining feature, is driving fairly reliable if incomplete dictionary services everyone’s way. Everyone?

The great oceanic divide is in fact just one easy to see version of a million micro barriers in a global media nexus. Readers, film and TV viewers also have to negotiate the shoals of accents and dialects inside each country. Local speech usage will be a permanent headache to understanding until we get real-time close captioning/subtitling for the less known lingos.

However, some bright spark has decided to provide an automatic ‘dialect translator’ for localizing your standard English into Geordie or Cockney variants and vice versa. You could easily extend this sort of tool to cover the U.S. and a dozen or so varieties/dialects/accents or whatever you call them spanning Canada, the U.S., the UK, New Zealand, South Africa, India and the rest of the Commonwealth. And so on and on through every language family and its rebellious offspring.

Yet the “countries divided by a common tongue” theme is not merely anecdotal. English in a variety of mutually understandable forms underpins much of global communication today, yet no single speaker of English as a 2nd language will learn in depth both US and UK (or other) local versions of every dialect or jargon, from slang and the highway code to legalese and pillow talk. So there will always be a need for tools that extract and manage appropriate glossaries on all these linguistic differences, and make them easily available to the casual or concerned user.

The range is growing: SMS messaging translators (from code to plaintext) will presumably be the next major case for treatment, as text messaging gradually takes over from voice mails as the preferred mode of asynchronous communication, at least among the now gen in Korea. If proved true, then messages in all languages will eventually require the full panoply of intelligent search, translation and other solutions so that message databases can play a role in the great digital conversation.

Quechua and Mapudungun go digital

0

Indigenous South American language communities are getting localized software from Microsoft. The Quechua version of Office XP and Office 2003 is coing out next year for the nearly 10 million Quechuaphones living across six or so countries. More intriguing is the project to localize into Mapudungun, the language of the Mapuches ethnic group (8% of Chile’s 15 million inhabitants). Come on OpenOffice

Cross-lingual search engine

0

Need to search the web with your language as the search input and another language for the result? Check out Babelplex, where this link shows the twin screen results for “machine translation” in English and French. It’s very fast, and offers the bonus of learning about comparative hits across languages (in this case, well over 6 million hits in English, against only 347,000 in French). Pity about that loathsome Babel morpheme doing overtime again.

Just nine

0

No one will have missed the ubiquitous English-language response to news headlining the blogosphere this week that Google’s eblogger service is being localized into nine languages. Wow!

What this actually means is that the service’s sign-in and account pages, are now available in Japanese, Traditional Chinese, Simplified Chinese, Korean, French, Italian, Spanish, German and Brazilian Portuguese. The next step will be to localize the pages containing Homepage, Login, Product Tour, Account Creation, Blog Creation, User Dashboard, and Edit User Profile.

It’s good news to potential bloggers who want to self-express in these languages, but not good enough for those speaking hundreds of other languages who want a simple online environment to write to and share comments from their communities or the world at large. What most astonishes me is how long it has taken to get this far. For an organization as powerful and rich as Google, surely the localization of an outside maximum of 5,000 words into 9 languages after “several years” of fielding its English language blog couldn’t be that hard or pricey.

As a global, language-aware technology provider, Google would have internationalized its basic blogger software right at start-up. And it has almost certainly anticipated the possibility of providing language-sourced advertising to the next generation of bloggers. Using the Yahoo groups model, for example, you might agree to finance your blog by allowing advertisers to target it automatically on the basis of the most salient words/ideas found in the blog. By opening up your language spread nine-fold (and potential punters several million fold), you then boost your advertising base. Not everyone would choose this option, of course, but since blogs are a platform for people with large mouths/egos, as well as for those virtuous few who simply want to share ideas, Google would realize that any way of attracting more readers into a given community is the name of the game.

So with the software ready and messages long stabilized, why wait? Google could even have tried tapping the enthusiasm of those who wanted blogs in ‘minority languages’ to get the job done. Think of the success of grass-roots online movements such as the Wikipedia (50 languages and rising) and indymedia in localizing reams of content fairly efficiently.

The really interesting task, though, is still ahead. According to the inexorable logic of the web, Google will eventually need to provide either a translation service for these blogs, or at least a search engine that can handle cross-lingual searching to let the linguistically challenged log the blogs. Most blog translations would probably not be worth the effort, but it would always be good for news junkies to know when someone has written something about a topic close to their heart. Just in case…

However, I presume that the ultimate market for cross-lingual access to the blogosphere might not be us Juans and Marikas of the web, but large-brand enterprises or governments snooping around to find out what’s being said about them among the unwashed masses, or trying to capture insights into ‘the next big thing’ in anything from voting trends to consumer tastes. Never forget that a technology that gives you the freedom to express your opinion can also reduce you to a small statistic in a database that others can use for their own ends.

Writing and reading content c. 2050

0

Most authoring done on personal computers today is just automated paper document creation or automated letters for an automated post office. Worse, the advent of WWW browsers for the Internet disastrously turned users into simple consumers, because the browsers do not permit any kind of balanced authoring for Internet content. That this was allowed to happen is almost beyond belief and has been a terrible setback.Alan Kay

What follows is a longish riff on imaginary authoring systems. Here’s the agenda: Could the steady shift from static text to dynamic visualization now emerging in computer programming, search results in business intelligence applications, product design and business process representations eventually impact the document authoring process? Should we start envisioning the ancient labor of writing as something that can be automated through diagrams and pictures? And could we later make it available to readers as pictures or movie narratives rather than text?

Consider this: instead of writing a document section by linear section, how about designing a document as a ragbag of ideas, examples, references, and stories, that can be automatically shaped and expressed by software, drawing on content databases and programming routines?

You could select from a menu of objectives, target audiences, etc (we’ll need a very powerful taxonomy of document types, functions, approaches, dimensions etc etc) and the machinery will do all the work. Stylometric tracking of your existing authored documents could produce a style profile that would be plugged in to personalize the output. But the ultimate purpose of authoring systems would be to craft content for readers, not ‘express’ the writer. Le style, c’est l’homme-qui-lit même….

This might suggest ‘Raymond Chandler does plasma physics’ or a legal tome à la Salman Rushdie, but the idea would be to explore how far personalization can be taken beyond its current rather banal horizon. Readerly authoring like this would truly herald that death-of-the-author meme, once identified by Barthes and explained by Foucault!!

Next logical step: instead of just reading a document in the usual old way, why not automatically translate it into a series of visual representations (a sort of Flash implementation) that do such things such as highlight arguments, capture inconsistencies, pick out story-lines, unveil subtexts, identify allusions, summarize, offer counter-evidence or arguments, and so on?

Rather as we entertain the notion of textual content as a sort of virtual microworld populated with concepts and arguments and stories that tug at our hearts and minds, let’s get the machines to automatically adapt this content into a physical movie, slide show, or flow chart, or indeed invent some new medium for the dynamic expression or represntation of ideational content, that simply does the job even better than scanning print with a pencil in one’s hand.

It is not hard to recognize the kind of documents that would lend themselves to this type of translation, starting with instructional texts. The history of culture is already packed with examples of ‘media translations’ from stories and jokes to plays to films to operas to comic books to multimedia extravaganzas to radio dramas to bowdlerized editions or signed versions. Think in English of the destiny of a piece of print content such as “Alice in Wonderland”.  But in a print culture, documents were scarce and were therefore designed to last. In a digital world, this is no longer necessary. We can now leverage the extraordinary capacity for media metamorphosis into a natural value of content holders. This rich personalization (way beyond mainstream genre or media metamorphosis) of textual content would therefore be one vital area to explore.

In other words, let’s start thinking of future authors’ words as (among other things) instructions to visualize (not imagine) content externally in the real world (i.e. not in our heads). And at the same time, let’s try and think of existing authored products as instructions to a machine to make visual or multimedia displays of both conceptual and intellectual content (“It is a truth universally acknowledged, that a single man is in possession of a good fortune, must be in want of a wife”), as well as more obvious descriptions or activities identified in texts by words (“The Marquis went out at five o’clock” or “the planet Mars”).

These automated ‘authoring’ and ‘reading’ (for lack of better terms) models would in the end converge into a new vision of content. Worthwhile documents get generated automatically from a resource by specifying a purpose (e.g. write a report summarizing activities in the fishing boat construction market in Norway between 200X and 200Y) and an audience (French bankers) and the software does the rest. The report’s “text” (as we would call it still today) could then be fed into a display model, offering a number of visualizations, including good old back print on white paper, but extending in various ways into dynamic representations that draw on real time content updates, group contributions etc etc.

None of these manifestations of a document needs a long shelf life. Indeed, ‘documents’ would become but fleeting coagulations of content in the constant flow of data, rather like the motile thoughts that stream through our consciousness. Actionable content alone would congeal into sets of instructions to act (strategies, tactics, plans, etc.).

To convince you that all this is not completely off the wall, here a few leads on concepts at work in the current content visualization space (see my earlier post on this topic). You are probably familiar with visualization desktop applications that can take a boring table of numerical data and turns it into colorful pie-charts and other graphics. Well, this track of presenting data as diagrams and other visual models has been pursued much further by firms such as Anacubis that helps analysts ‘discover’ knowledge by presenting reams of content as visually engaging diagrams that highlight links and associations that get lost in the gray blur of print. Another innovator here is i2’s investigative analysis software, that will automatically translate complex time line data into visual maps that help analysts compare the causes and effects of incidents, for example. In a similar way, some search engines will now offer pictures of results, representing the usual list of URLs as trees with branches, or star-shaped clusters of links.

This type of visualization is fine for inspecting an information source domain, but it is hard to know whether you could use it to picture the meaning (or implications) of a single document. Most of us still prefer the good old Table of Contents or the index to get a intellectual handle on texts. For example, anyone who has translated a large factual book will tell you that the best place to start is the index: translate the indexed terms and you have a powerful cognitive picture of what’s in the body of the book, plus a handy terminology base. My question here would be: are there any ways in which technology can help us ‘learn’ what is in a book (maybe I mean ‘read’ it) by having the book content engineered into a graphical and above all dynamic mode?

Another visualization track that gets us a bit closer to our initial visual authoring idea covers new products coming on the market designed to help word-centric people put their ideas into visual form to drive a product design process.

N8 Systems lets business analysts describe the steps of a business process in words and then automatically converts the language into diagrams.

The N8 text modeling tool then transforms the written requirements into use case and activity diagrams. The analyst is able to see immediately where the inconsistencies in the definition of a process or workflow occur, and can then make rapid iterations to achieve and articulate the desired outcome. He can then share the process diagrams with multiple constituents in the business unit to check that the definitions are accurate, and can quickly modify them as needed.

Once satisfied with the results, the business analyst provides the diagrams to the system architect, consultant, or other solution provider to help jump-start the requirements process. Communication is enhanced as both parties share precise, accurate diagrams of the written requirements.

Stottler Henke’s SimBionic is a visual authoring tool and runtime engine for creating complex behaviors found in games and instructional media.

It uses a graphical interface to specify behaviors, so that non-programmer ‘subject matter experts’ can create them. This reduces the risk of simulation errors related to miscommunication of content between an expert and a programmer.3D software

Perhaps the best example of the power of visualization to boost authoring is found in the emerging field of Product Lifecycle Management (PLM), where the French company Dassault Systems (DS) is leading the pack with the concept of 3D for All .

I happen know a little about DS because I once did some writing work for them. To simplify drastically, PLM software uses 3D functionality to model, design, engineer and test not just a potentially manufacturable product such as a mobile phone or an airplane door, but every other aspect of that product’s lifetime, from designing the appropriate manufacturing process to the kind of factory you’d need to house it in. In other words, rather as I was imagining for texts, products become huge databases of design information that can then intercat with any number of other digital tools. The key advantage in PLM is that you can work together as a team to deszign a new product and then test it virtually in 3D, instead of wasting good money on building a whole airplane in the real world to see what happens in the wind tunnel.

But what interested me most at DS was the disruptive fact that the lowly process I was involved in used none of this superb 3D software to enhance the process of producing a complex document. As in most text authoring situations, I would guess, it was all faxes, text files and PDFs of the graphic design and endless rewritings and no real integration between designer, graphics/colors/layout, and the words. Although I notice that DS has just signed a wide-ranging alliance with Microsoft to extend PLM software to small enterprise users of the Microsoft .NET platform, I doubt that document production will enter the PLM engineering mindset any time yet. I would, however, bet that games, advanced toys, and possibly business process design will be the first future targets for 3D – and other design – software.

Fractal power play at Spanish Language Congress

0

While eminent speakers of castellano discussed Linguistic Identity and Globalization’ at the Third International Spanish Language Congress in Rosario, Argentina, an alternative movement held its I Congreso de laS LenguaS (S for plurality, intercultura and metissage) in protest at their exclusion from the purely Hispano-centric discussions. At the 3rd Congress (held under the auspices of the Royal Spanish Academy, the Federation of Academies of the Spanish Language, and the Cervantes Institute of Spain ) they discussed ways of fighting the global dominion of English; the 1st CongresS protested through street theatre and Social Forum type discussions against Spain’s global Castilian thrall over other languages in the Iberian peninsula (Catalan, Basque and Galician) and above all in Latin America (Guranai, Quechua, Aymara and Qom). One man’s global is another man’s local. In the interests of fair reporting, here is part of the 1st CongresS manifesto expressed in the language of the 3rd International Spanish Language Congress:

¿Qué es realmente el congreso de la lengua? Un espacio cerrado donde todos somos excluidos.

No están invitados los docentes, los primeros trabajadores de la lengua y el habla.

Tampoco los estudiantes de todos los niveles, (incluso los universitarios especializados) y la mayoría de los intelectuales (lingüistas, antropólogos, escritores, etc.).

Los pueblos originarios sólo estarán “representados” por especialistas hispano parlantes que hablarán por ellos

En realidad, está excluido el pueblo. A todos los sectores antes citados se nos condena a efectuar “acciones paralelas” en el “marco del Congreso de la Lengua” y por fuera de él.

Es antidemocrático porque sus objetivos son injustos. Y por eso han decidido evitar la polémica.

Es que pretenden negar la realidad y la historia.

Niegan la realidad cuando parten de la ficción de que la lengua “castellana” es LA LENGUA de España y de la mal llamada “Hispanoamérica”.

Pero el español no es siquiera la lengua de toda España, allí subsisten, a pesar de haber sido reprimidas por siglos, otras lenguas como la vasca y la catalana.

Menos aún en América donde los españoles no pudieron nunca erradicar lenguas habladas hoy por millones como el guaraní, el quechua, el aymara, el qom; lenguas que en algunos países como Paraguay, Bolivia o Perú son mayoritarias.

Niegan esta realidad para avanzar en la imposición del “castellano” -en realidad del “castellano de la Academia”- y perseguir a los otros idiomas y “hablas” como “inferiores”, “menores”, “incultos”. Esto va a contramano de la necesidad de impulsar la alfabetización, la comunicación, respetando las distintas lenguas, las distintas formas del “habla” en sociedades divididas, sometidas y atravesadas por crisis profundas como las nuestras. Y donde los que menos hablan el “español de las Academias” son justamente los sectores más oprimidos, humillados y con difícil o nulo acceso a la educación formal.

En segundo lugar desconoce la historia, porque trata de borrar que la imposición de la lengua castellana es parte de una larga lucha de opresión en España y en América que incluyó la persecución y eliminación física de los árabes y judíos en España

En América la dominación lingüística fue parte indispensable del genocidio de los pueblos originarios; genocidio al servicio de uno de los más grandes saqueos que conoce la historia. Si no llegaron más lejos fue por la lucha de resistencia de esas naciones, luchas que forman parte indisoluble de nuestra identidad como pueblos y naciones.

Translation market in China

2

News just out from the Translators Association of China (TAC): China’s translation industry generated about 11 billion Yuan (US$ 1.32 billion) in 2003, and is expected to grow to over 20 billion Yuan (US$ 2.5 billion) in 2005. Or if you prefer this source, they even forecast that this market could double in the “not-too-distant future.” In other words, during the 2008 Beijing Olympics and 2010 Shanghai World Expo. The effect of large international events (e.g. Tokyo in 1964, and Seoul in 1988) on kick-starting the translation/interpretation professions in Asia is well known. And China’s entry into the WTO is obviously growing Chinese translation volume (and hopefully quality). However, the large current supplier base is considered to be poorly trained, structured and supported.

In November last year, the State Administration of Quality Supervision, Inspection and Quarantine issued an official Specification for Translation Service, which began in June, to provide objective criteria on translation qualifications and compulsory regulations. A certified translator examination system, China Aptitude Test for Translators and Interpreters, was also introduced last year. So far, about 30 percent of the 4,600 people who have sat the exam have passed.

I imagine this training effort is largely due to Wusun Lin, the then VP of the TAC, whom I met in Shanghai in 1998. At the time he told me that, there were almost no translator training programs worth the name, but estimated that there was a translator/ interpreter population of 500,000(!), 10% of which were “accredited professionals”. This figure of course covered translators working between the 55 official ethnic and linguistic minorities in the country. I notice there same figure is being quoted today, yet there is still “a 90 percent shortage in the number of qualified Chinese-foreign-language translators.”

The TAC claims there are currently over 3,000 translation companies operating in China. The number may actually be closer to 10,000 in that many small companies, which are registered as consultant agencies, actually conduct translation businesses. And they are not yet tooled up, though Trados, SDL and other extrenal tools suppliers are gradually making inroads. The People’s Daily reports that research on translation software in China was launched in the mid-1990s, but there are “only” 10 established domestic translation software companies, including China National Computer Software and Technology Service Co , Huajian Machine Translation Co and SJTU Sunway Information Technology Co., whose products are not really market-ready. In terms of population to translation technology company ratios, ten companies looks like a low figure. But is the ratio of active, sustainable translation technologies companies to populations any higher on a global level? The TAC rightly believes that as the market grows more competitive, translators will need to “look to technology to improve efficiency and quality.” But it is very unlikely that a local firm will be automating translation of the Beijing Olympics.

What’s missing from these reports is data on which languages are involved. We know about the need for Chinese localization of English and to a lesser extent European language content. We know about the need for translation into Simplified and Mandarin versions of Chinese. But it would be useful to know about volumes involving Chinese and other (Asian?) languages.

Slav(e) labor?

0

Backbytes reports on a job ad in the U.K. that went as follows:

European Systems IT and business development manager, Wimbledon…An Accountancy company based in London is looking for an All Rounder with the following skills: Visual Basic, Access, VBA, SQL, HTML and ASP. You must also have a Degree in IT or Computing and speak and write ALL of the following: English, Russian, Ukrainian, Bulgarian and Polish. Not too much to ask!’ For this, you can expect to earn £27,500 per year.

Global language stats for the web

0

Was a time when I used to visit Bill Dunlap’s pioneering Global Reach site regularly to find out about the fast changing figures on the language presence on the Internet. It doesn’t seem to feature any longer in data sources I come across, but someone’s blog prompted me to check out the latest (September) figures. Global Reach provides two essential data sets: year on year historical data to demonstrate evolution over time, and the sources for its figures.

English web users now account for 35% of the 801 million Web users, with Chinese users next with 13.7, trailed by Spanish (9%) and Japanese (8.4). The obvious figure, of course, is the forecast 100% leap for the online Chinese between 2004 and 2005, after already growing faster in recent years than any language constituency. See here for more details on the sources of these figures (in French).

People once used Global Reach online language data to demonstrate the need for localized e-commerce websites in a global marketplace. Now that this message has been widely taken on board, there is a need for more detailed stats that drill down inside these broad brush-stroke figures. I see a need for at least two critical sets of global stats, and I imagine there are technologies out there that can track this stuff automatically:

· the language spread found in globalized websites

· online language populations vs. country populations (to get a snapshot of movements in minority or immigrant language groups. E.g. see here for a report on the relevance of Latino language marketing in the U.S.)

Both these data sets should be visualizable synchronically and diachronically, so we can find out how fast web localization is proceeding, where it’s happening, and which languages are joining the localization pack. And governments as well as e-marketers will need national stats to inform educational, e-government and other policies.

If this is already being done, let’s hear more about it.