Battles won on the playing field

0

Scrutinizers of the spread of international English may like to note the two part series on “The Ascent of English” in The Financial Times published today (German managers on the value of English) and tomorrow (on English usage in China). You need to subscribe or take a free trial. Advice given to aspiring Anglophones includes:

Encourage employees from different countries to mix through an inter-company soccer league, which will lead to the use of many common and useful English words throughout the company.

I would have thought that, as in most contact sports, soccer players, as opposed to the game’s commentators, mostly bad-mouth each other on the pitch rather than talk about tactics, goals, teams and winning strategies. But though hard to get right in a foreign tongue, collective insulting is probably all part of the corporate game too.

The ethics of grammar checkers?

0

Sandeep Krishnamurthy, associate professor of marketing and e-commerce at the University of Washington, has written a withering critique of the Word grammar checker as a student proofing aid:

I knew Microsoft Word’s “Spelling and Grammar Check” feature was bad.  However, I never realized how bad this feature really was until a student turned in a poorly written report that was “spellchecked” and “grammarchecked”.  I have since tested this feature out hundreds of times.  My conclusion is that the “Spelling and Grammar Check” feature on Microsoft Word is extraordinarily bad (especially the Grammar check part).  It is so bad that I am surprised that it is even being offered and I question the ethics of including a feature that is this bad on a product that is so widely used.

Apparently MS has retorted in a subscription-only publication that the checker is designed “to catch the kinds of errors that ordinary users make in normal writing situations.” Seems to me that any ethically coherent training of students to use writing tools should include a discussion of what a grammar checker can and cannot do. There’s plenty of material on which to show where and when a checker is handy aid rather than an accomplished sub-editor. But a bored student in a hurry will always use anything that seems to offer a necessary minimum of quality. Methinks Sandeep doth protest too much. 

Transblogating

0

English first, then shock & horror: other languages don’t fit. Develop a technical fix. Which turns into a business opportunity. Until the next big digital thing starts the cycle all over again.

Sound familiar? Two decades ago, English-only software suddenly had to adapt to other language communities. Ohmygod! Double byte languages, different sort systems, tangible otherness. The technical fix was “localization”, which eventually evolved into the GILT industry and opened up a business opportunity for translators and their minders.

Just one decade ago, bright new web entrepreneurs suddenly realized you could reach further, inform better and sell more if you “globalized” your site into other languages. Ohmygod! Double byte languages, different sort systems, tangible otherness, lotsa money. The technical fix involved investing in globalization content management systems and streamlining workflows. And opened up a further opportunity for web site translators and their (richer) minders.

Now today, the rise and rise of blogs as the new personal publishing platform appears to be shuffling through the same old digital content choreography. First the new practice spreads through the English speaking/writing community. Then it’s rapidly adopted by other linguistic communities, probably starting with semi-bilinguals. The original content stream suddenly goes opaque. Ohmygod! We English-speaking first adopters can’t read Arabic, Pashto, Russian, Chinese… bloggers who have something important to say to us, given the ambient madness. Then comes the technical fix…

In fact we’re still looking for a technical fix, well exemplified in Tim Oren’s stream of ideas and proposals, and partially summarized in this recent Salon article.

So the longer term question is: will translating the blogosphere open up a new business opportunity? At first sight, no. Blogging today is personal, informal and anti-commercial. So friends or bloggers themselves are more likely to do the translating, if any. Yet witness the discussion about what’s happening to Joi Ito, a bloggerato who writes (in English) about digital copyright issues. It turns out that NEC is now re-purposing content from his blog, and translating it (professionally?) into Japanese.

Yet as the blogosphere grows, it will presumably develop tools that will trawl premium content from the oceans of self-serving static, and invent an appropriate business model to sustain that rich content stream, social software notwithstanding. So there is a strong chance that someone will seize on this new opportunity offered by this latest manifestation of Babel’s curse, and end this particular ‘new content’ cycle by delivering a cheap but effective transblogation service, based on whatever technical fixes make most sense.

Multilingual definitions on Google

3

Google will now do define: operations on words in other languages than English. As an example, check this result for “rough” which has various meanings in other languages than English, especially as part of the international vocabulary of golf. However, as one of the simplified Chinese sites returned shows, you don’t always get the results in the presumed language of the site – it could be an English word list hosted on that site. Thanks to Danny “The Search Engine” Sullivan for this news.

Cyril’s passport

1

In the same On Language column, William Safire cryptically ends on considerations about passports, computers and internet.

After mentioning current discussions on transliteration for administrative purposes, he claims that


[m]eanwhile, acting unilaterally, the Russian government has worked out its own plan for handling Russian names on its passports to make life simpler for immigration officials of other nations.

As far as I’m aware, Russian travel passports have used the French transliteration system for a long time, and I could not find any information about any current Russian plans to do anything about that. Can anyone enlighten us here?

Attempting an analysis, the column then states that the problem is that


most computer operating systems are based on the Roman alphabet. Maybe, like a new Caesar, the imperial computer will impose our present system on the rest of the world, forcing Slavic and Asian systems into our alphabet soup. Or maybe the United Nations will find a new raison d’être (that’s ray-ZON DET-ra) in standardizing a system to encode Roman and Cyrillic letters and Chinese and Japanese characters to make them computer-friendly on all the world’s screens.

And further:

For users of tomorrow’s Internet to accurately cross cultures, experts in phonetics and transliteration will first have to create and agree on a standard system.

Are we, the people of GILT, failing in our mission to such an extent that there are still people out there, writing columns in some of the most influential newspapers in the world, who think that computers and the Internet can only work with roman alphabets?

Nouillorque taillemse donne lailleque ze frentche?

0

A very strange On Language column in the NY Times about transliteration. After explaining why the actual phonetics of russian president Putin‘s surname do not transliterate very well into english, in particular because of the inappropriate hard t and in ending, William Safire weirdly goes on to

the reason that French is known as the language of diplomacy. In France’s official documents, as well as uniformly in the French press, Vladimir Putin’s last name is spelled Poutine. As a natural result, it is pronounced poo-TEEN, rhyming with our ‘’routine.’’ The French undoubtedly know that is not the way he or his compatriots, or even President Bush looking into his soul, pronounce Putin’s name.

The article omits to point out the French origin of its example word routine, and that it is therefore not pronouced, in original French, the same way as it is in English. In fact, not to oversimplify the point, the ine ending is not a lot like the English een, the latter being a long sound, whereas the former is a shorter one; it’s in fact somewhere between the English een and in.

One thing phonetology amateurs are familiar with, is that the soft t does exist in French, albeit not as an identified alphabetical or typographical marker: it softens automatically before certain sounds, such as (surprise!) i. If Russian has softener and hardener markers, it’s because they have both options in their phonetic structure. If French doesn’t it’s because it’s a direct consequence of the succession of the sounds.

Then there is the question of accentuation (stress). It is true that Russian is a heavily accentuated language, to the point that the location of the phonetic emphasis in the word will in many cases change the sound of the vowel. The o will be pronouced more or less as an english o if accentuated, whereas it will be closer to an ah when the accent is elsewhere in the word. Other vowels behave similarly. So accent is a key component of pronunciation in Russian, more so than in many languages. From that point of view, English’s natural accentuation on the first syllable is indeed closer to the original. But that’s not a function of a choice in transliteration spelling, it’s a function of the language’s stress rules.

So to summarize:

  1. The English transliteration gets the stress right, but is very approximative in the sounds;
  2. The French transliteration is better at getting the sounds right, but gets the stress wrong.
  3. In both cases, it is simply a direct result of each language’s phonetic rules.

The article continues the fallacious charge with:

Why the error in transliteration? Official French sources tell me that because the sound that we write as in has no place in French pronunciation, an e has been added to make the sound more amenable to the French tongue, and that’s all there is to it.

The author makes it clear that he considers this a false justification. One can only wonder about the familiarity of the writer with such questions: the e has not been added in French to make it more amenable, it’s been added because it would otherwise produce a different sound. It’s precisely adding the e that brings it closer to the English in. Mr. Safire seems to know that, because he then expands on how it would have been pronounced similarly to putain, meaning whore.

After failing to prove the point that purports to be the subject of the article, the article concludes by stating that the French

have embraced phony phonetics, unanimously choosing to mispronounce the name of the president of Russia

to avoid embarrassment.

It’s hard to tell which is the most important of the intentions in the article, between having a boyish chuckle on the French pronunciation of the English transliteration of Putin meaning whore in French, or portraying the French as a cowardly race who will go to any lengths, including debasing their own language, to avoid embarrassing either themselves or a ruler of Putin’s ilk. Either way, and notwithstanding Mr. Safire’s constitutional right to indulge in either, it would probably help language and culture professionals if such false ideas and wrong ways of thinking about language were not endorsed in such major public organs as the NYT.

Of course, the NYT’s job is not to make ours easier.

Localizing the business blogosphere

0

It is now banal to remind that never before in the history of mankind so much text and other content has been produced. Blogging has of course contributed to that explosion of text. Companies’ communication activities are also a big driver of this inflation, as well as consumers’ thirst for things to do when there’s nothing to do. We have entered a content-consumer-society, where content is a commodity similar to what consumer goods became a few decades ago. The difference is that, unlike for picnic forks or lightbulbs, anyone can produce content. So we now have however million blogs out there, producing free content.

So a new standard question now has arisen: how will all this content be translated? The argument goes that all this blogging thing is all about the flow real-time, always-on, free information and ideas, and that language is the last barrier to a truly free, flowing, global, real-time world of sharing, before we can all be brothers. As often in such utopias, it is a technological problem that prevents this from happening: so machine translation is the answer and it will work, because it has to work, because everybody talking to everybody at the same time in all the world and understanding each other is our collective destiny.

What about company content? Companies, another banality, need to speak to consumers in their language, and pull their cultural triggers, to sell or effectively engage with them. The kind of trust necessary to put someone else clothes or food on or in our bodies, to give one’s scarce money in exchange for an appliance, or to entrust our lives to their products, is elicited by such things.So, regardless of the now-boring claim by some that ‘100% publication quality’ (I’m not sure how this is measured) machine tanslation is for sometime in the next five to ten years, it seems unlikely that company content destined to engage people will rely on machine translation anytime in the foreseeable future.

Moreover we are to believe, if interested pundits are to be believed, that business blogging is the next big thing. The mere fact that you’re reading this blog should rather comfort that idea. At the same time, the question of the language of blogging, the tone, this new mix of familiar, personal yet well-crafted disposable written language, creating a deeper, more intimate relationship with the reader or audience, is the new hot topic. David Crystal even claims that it is an evolutionary event in the english language (there’s no reason to believe that this should not be the case in other languages, although the widespread appropriation of english by non-native speakers is probably a special case).

So how can the “real-time machine-translation” and the “personal tone, quality intimate relationship” approaches be reconciled? Well, they can’t, and the localization model for company blogging is certainly up for grabs. Should it be handled as any other marketing content and just budgeted as more words to translate? If blogging is the quintessential tool for creating intimate, relevent, personal and interactive relationships with your audience base, then this seems rather out of the question. So what’ll it be?

There are really five kinds of company blogs, as far as I can tell: 1- “unofficial” but company-facilitated employee blogs, 2- executive blogs, 3- user-community blogs, 4- industry punditry and 5- marketing blogs. Of these, 1-, 3- and possibly 4- are either non-applicable or nothing new in terms of localization. But executive blogs and marketing blogs are something different. These are where individual voices, seeking to interact with or influence customers, consumers, prospects, vendors, competitors, the media, their staff, express on behalf of a company or product. There is familiar insider technobabble there, hip, up-to-the-minute stuff, trackbacks, comments, and, not negligibly, a quagmire of implicitly shared cultural assumptions. Most effective communication, we know that by now, works mostly through such implicitness.

Running contrary to the dream of always-on pan-language communication through MT, such relationship tools cannot be globalized by other means than finding a voice in each market (or at least cultural and language area) that will be the catalyst of engagement. This can be the local MD, or a contracted local journalist or knowledgeable pro, or the marketing manager. Or, more simply and probably more appropriately, a local company team-member with the right attitude.

Because that’s the thing: it’s not so much about pushing messages, it’s about having a voice. Can you localize a voice?

Doug Piranha software

2

Can you automatically detect affect in linguistic media such as speech and text? In December, I mentioned a Scottish company that claims to identify mood in voice and use the knowledge to power a safer driving system. Spotting affect in text as an indication of a product/company/brand/policy/person’s reputation, say, is possibly harder. You can’t just go for keywords such as “bad” or “brilliant” in case they’re embedded in a negative expression, as the New Scientist kindly explains in a news item on UK company Corpora Software.

“Corpora has come up with a program called Sentiment, which uses algorithms to tease out grammatical components, such as nouns, verbs and adjectives, and identify the subjects and objects of verbs. It can even analyze pronouns like “it”, “he” and “her” to work out what words or concepts they are referring to.

Having an understanding of grammatical structure makes it possible to filter out words that are not relevant to the sentiment of the article, Jacobi says. So instead of assuming certain words, such as “unpredictable” or “rubbish”, are positive or negative it allows the structural context to disambiguate them.”

When the web first came along, everyone used purely formal indicators such as page hit numbers as a sign of reputation. Now we have entered the actual content, and use keywords in search engines that can dredge up an apparently relevant advertisement and place it on your page. Unsurprisingly, if you Google on “stupid” “hateful” and “rubbish”, you get no ads. Since humans can perform metalinguistic operations on a “bad” word to produce an ironical effect, we are now obliged to start parsing texts to see whether the grammar of words can automatically hint at attitude. Presumably, a first step towards some sort of semantic processing of linguistic viciousness. Another ten years ( as they proverbially say in the NLP industry) and we could be reading about a Piranha Brothers engine. Remember Monty Python’s Doug Piranha:

He used… sarcasm. He knew all the tricks, dramatic irony, metaphor, bathos, puns, parody, litotes and… satire. He was vicious.

Might have been a description of computer HAL in the 2001 movie?

Program writing on the wall?

0

For most of the past 4 millennia, writing (inscribing linguistic and related signs on a visible, semi-durable surface) has been an elite skill. Mass literacy only emerged in the 19th century in Europe, and is still on the global agenda of development organizations. You have to learn to write, and that takes time and resources. And ceteris paribus, some writing systems must be cognitively “easier/quicker” to learn than others.

In the past, scholars, monks and writing masters jealously preserved their magical writing skills, especially in cultures such as Sumer and Egypt, where the writing system was highly complex and required a lengthy education. You still find a specialist caste of public writers armed with typewriters lined up outside post offices in countries where illiteracy is common.

Writing computer programs is also an elite skill, and we depend largely on a caste of trained operators who jealously protect their magical ability to translate complex tasks into software for us. Question is: will programming, like writing, evolve into mass computeracy over the next decade/century/millennium? Will we eventually be able to write our own routines and use the inherent power of the machines themselves to transform task concepts expressed in natural language into executable software routines?

For the beginnings of an affirmative answer, check out this story about researchers at MIT who are experimenting with a program called

Metafor that provides a visualization framework for people to program “stories” using English. The idea is to leverage the expressive power of natural language syntax and a semantics as a programming language. Put crudely, you need a translator that transforms natural language to code. Where we might need help from the global developer community in the long term is in making sure that all natural languages, not just English, can be used to program code. There’s no point in attenuating the hold of one “closed shop” over our practices if all we do is replace it with another – this time linguistic – elite.