Friday, 31 December 2010

Evolving English

Yesterday I finally got round to visiting the Evolving English exhibition at the British Library. (The website also has an associated blog.)

The exhibition is excellent, well-organized and full of interest. Quite apart from the linguistic information it offers, it is fascinating to be able to see the actual Lindisfarne Gospels, an original manuscript of Beowulf, Caxton’s printing of the Canterbury Tales, a Shakespeare quarto, and — among more modern treasures — John Betjeman’s heavily revised manuscript of How to Get On in Society (‘Phone for the fish-knives, Norman’). David Crystal gives an introduction on a video loop. Over headphones I listened again to Kenneth Williams and Hugh Paddick’s polari spiel from Round the Horne.

Among the items of phonetic interest are a leaden Punch joke about h-dropping by the lower classes and an excerpt from the hard-to-believe 1929 recommendations of the BBC Advisory Committee on Spoken English.



The exhibition continues at the British Library until 3 April. Admission is free.
___

Happy New Year to all. This blog will be suspended during the whole of the month of January 2011. Next posting: 1 February.

Thursday, 30 December 2010

ban legacy fonts!

Do you remember the bad old days before Unicode? The time when there was no standardized way of encoding phonetic symbols? when word processing was single-byte and fonts were 8-bit, so that any given font was limited to under two hundred characters? when the various phonetic fonts available all used different encodings, so that where one person had input ɥ another might see ɦ or ʰ or something else entirely arbitrary? when if you transferred a document to a different computer you would as likely as not get garbage for your phonetic symbols? when your Powerpoint presentation using the computer supplied by local organizers would probably fail to display your phonetic symbols properly?

Thank goodness those days are past. Nowadays we all use Unicode, the internationally agreed industry-wide font-encoding standard for all alphabets and scripts, covering all the languages of the world as well as all the phonetic (and other) symbols we might need. A single font can now contain thousands, indeed tens of thousands, of different characters. So we no longer have to keep switching fonts merely in order to include phonetic symbols. In this blog I can be confident that when I input a particular phonetic symbol you will see that same phonetic symbol on your screen, no matter where you are and no matter what platform you are using. (OK, there may be marginal cases where the font you are using falls down over one or two unusual symbols: but then you will probably see a blank square or something similar — you won’t see the wrong phonetic symbol or some ludicrous webding, as used to happen.)

I celebrated this progress and documented the details in the poster paper I gave at the 2007 International Congress of Phonetic Sciences in Saarbrücken. (If you’re interested, here’s the printed version.)

But phoneticians haven’t all caught up.
The next ICPhS is due to be held in Hong Kong in a few months’ time. The deadline for paper submission is the beginning of March, so it’s time for everyone to get their thoughts in order and start writing. The Call for Papers page on the conference website gives the following instructions about phonetic symbols in submitted papers.
• One of the following IPA fonts is to be used for congress papers:
IPA-SAM phonetic fonts: http://www.phon.ucl.ac.uk/shop/fonts.php
SIL phonetic fonts: http://scripts.sil.org/encore-ipa-download

What are these fonts, so brusquely prescribed?
  • The IPA-SAM fonts are 8-bit fonts that I created around fifteen years ago. Building on SIL software, they enjoyed some considerable popularity because the encoding and therefore the keyboarding fitted in nicely with the way phoneticians actually use phonetic symbols. Nevertheless, once Unicode became available it rendered these and other specialist 8-bit fonts obsolete. For the last five years or more I have been actively discouraging people from using the fonts I created, because Unicode phonetic fonts are now widely available. Indeed, more and more of the ‘core’ fonts supplied with new computers include all the IPA symbols. So everyone should now use Unicode rather than ‘legacy’ fonts like the IPA-SAM fonts.

  • If you follow the ICPhS link to the SIL site, you will see this notice, prominently displayed.
    Important
    The SIL Encore IPA and SIL IPA93 fonts are obsolete, symbol-encoded fonts. Their use is discouraged. If you decide to download and use these fonts, please note there is no user support for these fonts.
    If your university or organization requires the use of these fonts, please request they change their requirement to Doulos SIL, a Unicode-encoded font which contains the complete IPA repertoire.

    Yes, their use is discouraged. Did you read that, conference organizers?

The Word template supplied by the organizers for ICPhS conference papers contains the following.
Phonetic fonts
You can use phonetic symbols and special characters in your paper. To make sure that readers of your article can see the phonetic symbols in the PDF document, all special symbols must be embedded in the PDF. Depending on the software you use to produce the PDF the details may vary. In our experience the fonts are usually embedded, but this can be checked e.g. by inspecting the "Document Properties -- Fonts" in Acrobat Reader.
It is recommended to use one of the following fonts to show phonetic symbols (links for free download can also be found at the Congress website):
• IPA-SAM phonetic fonts [3]
• SIL phonetic fonts [4] (Unicode is accepted)

“Unicode is accepted.” As an afterthought. Big deal.

Where have the congress organizers been for the last ten years? Unicode should be required. And legacy fonts firmly deprecated.

Wednesday, 29 December 2010

1949 revisited

The December 2010 issue of JIPA (Journal of the International Phonetic Association) celebrates its own forty years of publication and the 125th anniversary of the publishing of a journal by the IPA. Prior to 1970, the journal was known successively as Dhi Fonètik Tîtcer, dhə fənetik tîtcər, ðə fonetik tîtcər, lə mɛːtr fɔnetik and, for the seventy-five years from 1895, lə mɛːtrə fɔnetik.

To celebrate this anniversary, the current issue includes a complete scanned reproduction, with original cover pages and pagination, of the 1949 booklet Principles of the IPA. This booklet comprised (i) a theoretical introduction explaining the association’s alphabet and the principles for its use, and (ii) exemplification by phonetically transcribed texts in some fifty different languages. It is now accompanied by a short introduction written by Mike MacMahon, the IPA’s historian and archivist.

MacMahon mentions two misprints in the 1949 text, commenting that their appearance is not surprising “given the complexity of setting phonetic texts in a pre-computer age”. One is a missing diacritic. The other is the ‘problematic’ placing of a raising diacritic next to [u] (thus ) in the Afrikaans specimen. (Since cardinal u is by definition as close/high as is possible without crossing the vowel limit line into friction, it can hardly be raised further.)

What makes the latter more mysterious is that the same problematic combination is to be found in the specimens of Tswana and of Scottish Enɡlish — which MacMahon does not mention. Yet we know that Daniel Jones, the editor of the 1949 booklet, was careful to the point of obsessiveness about the exact typographical form of the phonetic symbols he used.

There are other misprints. In the specimen of Finnish we find riːsiu for riːsui, in that of “Roumanian” dz for , in the Welsh emlaen for əmlaen. There are doubtless others. Among factual deficits, the Japanese specimen lacks all mention of pitch accent.

Although it is not a misprint, it is shocking to find that as recently as half a century ago the name of the language Xhosa is supplemented by the now grossly offensive gloss “(Kaffir)”.

The English (“one variety of Southern British”) text of The north wind and the sun is transcribed, for illustrative purposes, in three different forms, one “broad”, one “slightly narrowed” and the third “still narrower”. This third transcription, reproduced below, contains two striking inconsistencies. One of the narrowings involves the explicit symbolization of schwa as opener (ɐ) in final position than elsewhere (ə). But if stronger is ˈstrɒŋɡɐ in the first line, why is it ˈstrɒŋɡə in the fourth? If that is a subtle observation of the effect of a close-knit following than, why, given ˈstrɒŋɡɐ and ˈtrævlɐ, is other in final position not ˈʌðɐ (line 4)? And why is the MOUTH vowel written in əˈraɷnd (line 6) but au in ˈaut (line 7)? Like Homer, DJ evidently sometimes nodded.


Tuesday, 28 December 2010

merry Mary and hairy Harry

An American academic, not a phonetician but working in a related field, sufficiently eminent that I have heard his name and even read one of his books, wrote to me asking for help in puzzling out the sets of words in which he, like many other Americans, makes no distinction.
I pronounce “merry”, “Mary”, and “marry” as homophones, as many Americans do, but a good many other Americans (and, I believe, a higher proportion of Brits) pronounce all three differently.

I referred him to vol. 1 of Accents of English, which he seems delighted with.

It’s not just that “a higher proportion” of Brits distinguish the three sets. As far as I know, all do. To the best of my knowledge no native speakers of English outside north America lack the three-way distinction merry — Mary — marry (RP ˈmeri, ˈmeəri, ˈmæri). We do not rhyme sharing with herring. We do not rhyme clarity with prosperity.

Just as this fact may come as a surprise to Americans, and seem problematic and mysterious, so it can be a surprise for non-Americans to find that some Americans make no distinction. And Americans can therefore get confused over spelling in cases where we never would.
Words like merry belong under DRESS, those like marry under TRAP, and those like Mary under SQUARE, which historically is derived from FACE. So you can often tell which word belongs where by reference to the spelling. Before double rr, the spelling e indicates DRESS and a indicates TRAP. With a before single r the vowel may be SQUARE, as in vary, parent, aware, compare, garish, Carey; but because our spelling system does not consistently distinguish short and long vowels before a single consonant letter it may also indicate TRAP, as in arid, apparent, comparison, circularity, Gareth, Gary. The suffix -ary is a special case.

The DRESS-TRAP distinction, as we know, can be a trap in EFL. A Danish friend of mine, now dead, was telling me about an acquaintance whose name I heard as Berry. An unusual name, I thought, but not impossible at least as a surname. It was only years later that I discovered he was really Barry.

The day before yesterday the Sunday Times travel section had an airline cabin attendant recounting how
a very anxious-looking couple boarded a Chicago–Heathrow flight. … They were studying the route maps intensely and staring wildly around the cabin. I asked if everything was all right, and the gentleman said, in a broad Midwestern American accent, ‘Are we all going to perish?’ I thought ‘Oh dear, here we go’, and assured him that we were not, but he just became even more upset, pointing at himself and his wife, then saying ‘We’re going to perish.’ I put my hand on his shoulder and told him in my most soothing voice that there was no way that they or anyone else on the plane was going to perish, but this had the reverse effect. It was only when my supervisor came over that we realised that they were going to Paris, and hadn’t realised they had to change planes at Heathrow.

Monday, 27 December 2010

countless thousands

As soon as I watched the brief preview of the Queen’s Christmas speech on Sky News I noticed her pronunciation of the phrase countless thousands (blog, 8 December).

I was not the only one. Edward Aveyard writes
I noticed your recent post on whether the Queen really uses /ai/ for MOUTH or not. If you listen from 3:40 to the Christmas message, she says “countless thousands”. I hear in countless but in thousands. … It's almost as if she had been reading your blog and wanted to give you something to analyse.

On Christmas Day I tried to download the video of the speech from the BBC website, but without success: although you could watch it you couldn’t save it. I looked on YouTube, but it wasn’t there. Now, though, TheRoyalChannel has uploaded it to YouTube: thanks, Edward, for the link.

Listen here, at 03:45, for the phrase in question.

I agree with Edward’s judgment. Watch HM’s lips in each of the two MOUTH tokens.

Another interesting pronunciation is powerful, here at 03:17.

It appears to be fully smoothed and compressed, ˈpaːfl̩. This is how I often pronounce that word myself, though some people seem disinclined to believe me when I assert that this reduction is widespread in RP. In my analysis, the “smoothing” process removes the second element of a diphthong, in this case MOUTH, when before another vowel (aʊ ə → a ə — or, of course, it could equally well have been aɨ ə → a ə). Then the “compression” process reduces two syllables to one (a.ə → aə). Finally, the monophthongization process suppresses the second element of the resulting diphthong, with compensatory lengthening of the first element (aə → aː). Thus a possible ˈpaʊ əf l̩ is reduced to ˈpaːf l̩. All three processes are variable (optional) and rule-governed (systematic).

There was a Two Ronnies sketch about misunderstandings arising from PRICE-MOUTH confusion (ground misheard as grind, etc). Can anyone locate it on YouTube or elsewhere?

Thursday, 23 December 2010

Denisovans

Today’s newspapers carry news, based on a report in the science journal Nature, of DNA findings relating to an archaic group of humans, some of whose fossilized remains have been found in the Altai mountains of southern Siberia. (Here’s the Guardian’s version. There’s also an informative article on the “Denisova hominin” in Wikipedia.)

The new human ancestors were named Denisovans after the Denisova cave in the Altai Mountains where their remains were found.

The matter was duly reported on BBC R4 in this morning’s Today programme.

But how do we pronounce Denisova and its derivative Denisovan? In particular, where does the stress go? The BBC reporter stressed the second syllable, -ˈnɪs-.

The name of the cave is of course Russian, and is written in Cyrillic as Денисова. It is the feminine of Денисов Denisov, from the name Денис Denis. But the stressing of Russian patronymics ending in -ов (-ov) is notoriously unpredictable.

None of the pronunciation dictionaries I have to hand record the name. But the online resource Forvo does!

(Forvo is a website with sound files demonstrating the pronunciation of a claimed 800-thousand-odd words in 267 languages. Anyone can upload a sound file showing how they pronounce a given name or word.)

A speaker described only as “Female from Russia” pronounces Денисов clearly as dʲɪˈnʲisəf. Isn’t the internet wonderful?

Assuming that this is the regular Russian pronunciation, it follows that Денисова Denisova is dʲɪˈnʲisəvə and that we should anglicize it as dəˈnɪsəvə (or perhaps with dɪ- or de-, or indeed with -ˈniːs-). The hominins, then, are dəˈnɪsəvənz.

This was indeed the pronunciation used by the BBC presenter. Well done the BBC Pronunciation Unit.
_ _ _

Happy Christmas to everyone. Next posting: 27 December.

Wednesday, 22 December 2010

a credulous scientific report

The December issue of Scientific American carries an article entitled ‘A Click of the Tongue: Ultrasound Translates Dying Languages’. (Thanks to the Clinical Linguistics blog for this link, supplied by Madalena Cruz-Ferreira.)

The article is about the use of ultrasound imaging to study articulation.
This portable technology, which became affordable to linguists around 2000, allows researchers to see the tongue as it moves in real time. It is one of the only medical scanning devices that can keep up with speech; MRIs, for example, are too slow.
Thanks to this emerging technology, [researchers] have documented some of the fastest sounds in human speech: the click consonants present in many rare African languages. Because linguists did not know exactly how the clicks were produced, the sound was placed in a “mixed-bag” category of the International Phonetic Alphabet.

Up to a point, Lord Copper. (OK, if you don't get the reference, go here.)

I wonder what evidence there is that clicks are “some of the fastest sounds in human speech”. Impressionistically, I’d have said that clicks (= sounds produced on a velaric ingressive airstream) are no faster or slower than sounds produced with pulmonic or glottalic airstream mechanisms. I suppose the claim is trivially true, in that the postalveolar (retroflex) click [!], for example, is similar in production to the plosive [ʈ], except that it involves a different airstream mechanism. And plosives are fast(ish). The dental click [|], on the other hand, is typically somewhat affricated, which means it is not so ‘fast’ a ‘sound’. And presumably pulmonic-airstream taps and flaps are the fastest of all.

It is not true that previously linguists “did not know exactly how the clicks were produced”. We can quibble about what knowing something “exactly” might be, and who the unspecified ‘linguists’ are; but phoneticians have been familiar with clicks and their production at least since the 1930s. There is a clear account of click production in, for example, Westermann and Ward’s 1933 Practical Phonetics for Students of African Languages. Their schematized diagrams are pretty good, too.
It was Pike, in his Phonetics (1943), who systematized the classification of airstream mechanisms, describing that of clicks as “velaric ingressive” (or “oral ingressive”).

Nor is it true that clicks are ‘placed in a “mixed-bag” category’ on the IPA Chart. They are in a box clearly labelled Consonants (non-pulmonic), along with the implosives and ejectives that occupy the other columns in the box.

Miller’s research, published in 2009, may well have “organized the clicks based on attributes such as airstream (where the air comes from), place (where the mouth constricts) and manner of articulation”. But in doing this she was merely following a long and established tradition. It is fifty years since I was taught this way of classifying clicks (and all other speech sounds), and I passed it on to my own students throughout my teaching career. I assume all teachers of general phonetics do likewise.

And of course ultrasound imaging in no way enables us to “translate dying languages”, though it might aid us in describing them.