skip to main |
skip to sidebar
Yesterday’s posting called for the small-cap-A symbol. I coded it straightforwardly in HTML as <small>A</small>. But blogspot accepts far fewer HTML tags in comments than it does in postings, so Paul, commenting, successfully entered it as a distinct Unicode entity, U+1D00.
Many, though by no means all, alphabetic small capitals are available in the Unicode range 1D00 to 1D7F. This block is known as Phonetic Extensions, and carries the introductory note These are non-IPA phonetic extensions, mostly for the Uralic Phonetic Alphabet (UPA).
The small capitals, superscript, and subscript forms are for phonetic representations where style variations are semantically important.
For general text, use regular Latin, Greek or Cyrillic letters with markup instead.
As well as small caps (ᴀ ᴁ ᴄ), superscripts (ᴬ ᴭ ᵃ) and a few subscripts (ᵢ ᵣ ᵤ), the block contains various other typographically interesting characters. (I have no idea what they are used for in the Uralic Phonetic Alphabet — though see here.)
Here among the small caps you will find a ‘reversed N’, ᴎ, a sideways Ø (ᴓ) and a sideways ü (ᴞ). There is a ‘Latin letter voiced laryngeal spirant’ (ᴤ) and a ‘Latin letter ain’ (ᴥ).
Not everything here is from the UPA. There is also a special ligature ᵫ, which I can see appealing to English lexicographers who prefer respelling to proper phonetic symbols, as will ‘Latin small letter th with strikethrough’, ᵺ. There is also something called ‘insular g’, ᵹ, labelled ‘older Irish phonetic notation’.
Although they are not official IPA symbols, users of IPA will be happy to find here the lax high vowel symbols ‘with stroke’, ᵻ ᵼ ᵾ ᵿ: two of these are used in the Oxford Dictionary of Pronunciation, though the first, ᵻ, bears the Unicode warning ‘used with different meanings by Americanists and Oxford dictionaries’.
A further Unicode block, Phonetic Extensions Supplement (1D80 to 1DBF) covers various former IPA symbols from which recognition was withdrawn at the Kiel Convention in 1989: those for consonants with velarization ᵬ ᵭ ᵮ ᵯ ᵰ ᵱ ᵲ ᵳ ᵴ ᵵ ᵶ and palatalization ᶀ ᶁ ᶂ ᶃ ᶄ ᶅ ᶆ ᶇ ᶈ ᶉ ᶊ ᶋ ᶌ ᶍ ᶎ, and for both vowels and consonants with retroflexion ᶏ ᶐ ᶑ ᶒ ᶓ ᶔ ᶕ ᶖ ᶗ ᶘ ᶙ ᶚ. So we can now find in Unicode everything we might need in order to digitize the 1949 IPA Principles, Jones’s The Phoneme, and various English-language accounts of Russian phonetics.
Commenting on Monday’s blog, Wojciech made the surprising remark Re the symbol 'a' in IPA: I too find it strange that it's reserved for a phoneme which occurs so rarely in European languages (if it occurs at all). Whereas the common continental (and Northern English, methinks) 'a' has got to be transcribed 'ä'.
I say no it isn’t, and no it doesn’t.
The vowel a occurs extremely commonly in European languages (and of course in non-European languages). The Northern English TRAP vowel, too, is very satisfactorily represented by the symbol a, with no diacritics. The contrary claims reveal a basic misunderstanding of how phonetic symbols are used when we represent the phonemes of a language or language variety. Let’s see why.
The symbol a is one of the set of symbols representing the ‘Cardinal Vowels’ i e ɛ a ɑ ɔ o u defined by Daniel Jones.
No language is actually spoken with cardinal vowels: they are idealized reference points not defined by what happens in any particular language. (They are, however, suspiciously similar to a subset of the vowels of standard French as spoken in Jones’s day — though the quality of French ɔ, at least, was and is considerably different from that of cardinal ɔ. In passing we may note that the articulatory-auditory theory behind Jones’s cardinal vowel scheme is no longer accepted.)
Rather, these symbols are used for vowels in the general area concerned. Like all IPA symbols, they allow some considerable leeway. A typical French e is not identical with a typical Italian e or a typical German e, although all share a general similarity and all can be characterized as unrounded, front, and close-mid (‘half-close’). Compare colour terms, where we happily refer to shades of crimson, scarlet, vermilion and so on all as ‘red’. We are dealing not with discrete entities but with points in a multidimensional continuum.
In those languages it so happens that the close-mid e is distinct from an open-mid (‘half-open’) ɛ. (This claim is subject to qualification: for many French speakers the choice of one or the other can be more or less predicted from the phonetic environment, although others distinguish e.g. les le from lait lɛ; not all Italians make the distinction between venti ‘twenty’ with e and venti ‘winds’ with ɛ; in German the vowel quality distinction is accompanied, in stressed syllables at least, by a length distinction.)
There are many other languages in which there is only one unrounded mid front vowel: they include Greek, Spanish, Serbian, and Japanese. Qualititatively this may lie anywhere between cardinal e and cardinal ɛ. In each case the appropriate symbol, though, is e. In the words of the 1949 IPA Principles booklet (§20), When a vowel is situated in an area designated by a non-roman letter, it is recommended that the nearest appropriate roman letter be substituted for it in ordinary broad transcriptions if that letter is not needed for any other purpose. For instance, if a language contains an ɛ but no e, it is recommended that the letter e be used to represent it. This is the case, for instance, in Japanese…
Similarly, the symbol a, which as a cardinal vowel symbol denotes an unrounded front open (low) vowel, is also appropriate to denote an unrounded open vowel of any degree of advancement (anywhere from fully ‘front’ to fully ‘back’) if that is the only open vowel in the language. This is the case in Spanish, Italian, Greek, Serbian, German, and Polish, to mention only a handful of European languages. It is also the case in thousands of other languages around the world.
In RP I say ðə kæt sæt ɒn ðə mæt. If I switch into northern (I was bidialectal as a child), I say ðə kat sat ɒnt mat. That’s how I would transcribe it. I’ll leave someone else to measure the formant values of my northern a to determine just how central it might be.
This is a live issue. The Council of the IPA, having previously failed to agree, is again debating the issue of whether to recognize an additional vowel symbol, A, to represent a quality between cardinals a and ɑ. I shall vote against.
The Guardian has a regular rubric in its Corrections and Clarifications column, Homophone Corner. Yesterday’s read as follows.
This led me to wonder what proportion of NSs have illicit (illegal) and elicit (evoke) as categorical homophones. Most of us, for sure. But are there some who make the vowel of elicit tenser than that of illicit? And do they do this variably or categorically?
I ask because this is relevant to the notation appropriate for the Latin prefix e- in English words. As you will be aware, for the third edition of LPD I simplified the notation for the unstressed prefixes be-, de-, pre-, re-, deciding to use the happY vowel i rather than enumerating mainstream ɪ plus variant iː. (In any case we still need the further variant with ə.) I really wasn’t sure whether to include the e- words in this, but eventually decided to.
Even that decision left marginal cases that were difficult to decide, and for which I may with hindsight have made the wrong decision. Elect? Event? Eleven? Of course, the decision for each particular word must depend not on etymology but on whether there appear to be people who use the tenser vowel — hence the inclusion of eleven, which does certainly not contain Latin e-.
It also means that the main pronunciation given for elicit, iˈlɪsɪt, looks different from that for its putative homophone illicit, ɪˈlɪsɪt, which clearly has no tense-vowel variant. (Compare the main prons for descent diˈsent and dissent dɪˈsent, which likewise are homophones for most speakers but I think not all.)
For previous discussion of the general issue, see my blog for 29 Jan 2007.
kɪdɪŋ ɔː nɒt, ɪf ə kɒmənteɪtər ɒn fraɪdiz blɒɡ kleɪmz tu əv hæd dɪfɪkl̩ti prəʊsesɪŋ ðə hedlaɪn ðen ɪts haɪ taɪm wi hæd ənʌðər entri rɪtn̩ həʊlli ɪn fənetɪk trænskrɪpʃn̩. (ðə lɑːs sʌtʃ entri ɪn ðɪs blɒɡ wəz ɪn dʒuːn.)
lɑːs naɪt aɪ pleɪd maɪ mələʊdiən ət ə seʃn̩ ɪn ə pʌb ɒn wɪmbl̩dən kɒmən, nɒt veri fɑː frəm weər aɪ lɪv. ðiːz seʃn̩z ə held wʌns ə mʌnθ ən ɔːɡənaɪzd baɪ ə ləʊkl̩ mɒrɪs saɪd.
dʒʌst ʌndə twenti piːpl̩ tɜːnd ʌp fə ðə seʃn̩. ðeɪ ɪŋkluːdɪd θriː ʌðə mələʊdiən pleɪəz. ɪts ɔːlwɪz ɪntrəstɪŋ tə kəmpeə nəʊts. bifɔː wi stɑːtɪd, wʌn əv ðəm kaɪndli əlaʊd mi tə traɪ aʊt hɪz ɪnstrəmənt (mʌtʃ mɔːr ɪkspensɪv ðəm maɪn).
evriwʌn wəz siːtɪd əraʊnd teɪbl̩z ɪn ə smɔːl rʊm ɪn ðə pʌb (ðə snʌɡ). wʌns ðə seʃn̩ prɒpə wəz ʌndə weɪ, ðə fɔːmən (tʃeəmən) kɔːld ɒn iːtʃ pɑːtɪsɪpənt ɪn tɜːn tə liːd ə tjuːn ɔːr ə sɒŋ. ðə prəʊɡræm wəz ə mɪkstʃər əv ɪnstrəmentl̩ stʌf (fɪdl̩z, kɒnsətiːnə, maʊθ ɔːɡən, mələʊdiənz) ənd ʌnəkʌmpənid sɪŋɪŋ. tuː ruːlz əplaɪd, əz ɪz juːʒuəl ɪn pʌb seʃn̩z — nəʊ æmplɪfɪkeɪʃn̩ ən nəʊ pleɪɪŋ ɔː sɪŋɪŋ frəm ə rɪtn̩ skɔː.
ðə sɪŋəz sæŋ veəriəs fəʊk sɒŋz ən fəʊk-staɪl sɒŋz. wiː ɪnstrəmentl̩ɪss pleɪd ɪŋɡlɪʃ (ənd ʌðə) dɑːns tjuːnz. ðiːz ə tɪpɪkli θɜːti tuː bɑː riːlz, dʒɪɡz, hɔːnpaɪps ɔː wɔːltsɪz, wɪð ðə strʌktʃər AABB. ðə kənvenʃn̩ ɪz ðət ju pleɪ iːtʃ tjuːn θriː taɪmz θruː, ɡɪvɪŋ ʌðə pleɪəz taɪm tə pɪk ʌp ðə melədi ən dʒɔɪn ɪn ɪf ðeɪ kæn.
maɪ əʊn fɜːs kɒntrɪbjuːʃn̩ wəz ə raʊdi riːl kɔːld tʃaɪniːz breɪkdaʊn (Chinese Breakdown), wɪtʃ tə maɪ səpraɪz ði ʌðə pleɪəz dɪdn̩t nəʊ — ɪt wəz wʌn əv ðə steɪpl̩z əv ðə bænd aɪ juːs tə pleɪ ɪn fɔːti jɪəz əɡəʊ — fɒləʊb baɪ ðə krʊkɪd stəʊvpaɪp (Crooked Stovepipe). leɪtə, wem maɪ tɜːn keɪm raʊnd əɡen, aɪ pleɪd dʒesiz hɔːnpaɪp (Jessie’s Hornpipe), seɡweɪɪŋ ɪntə səʊldʒəz dʒɔɪ (Soldier’s Joy), wɪtʃ evriwʌn nəʊz.
In the talk on Multicultural London English that I recently gave in Japan, one of the things I mentioned was a tendency to simplify the phonetics of the indefinite and definite articles by reducing their allomorphic variation. My data came from Kerswill et al., ‘Contact, the feature pool and the speech community: The emergence of Multicultural London English’, Journal of Sociolinguistics 15/2, 2011: 151–196.
I am well aware that MLE speakers are not the first NSs to fail to observe the rules that we teach EFL students for the pronunciation of the (that is, ðə before a consonant sound, ði in front of a vowel sound, plus the occasional strong form ðiː). Indeed, I make the point in the note I put in the relevant entry in LPD. 
What seems to be true is that ðə plus hard attack before a word beginning with a vowel sound is more frequently heard in MLE than in, say, traditional Cockney or RP. But this is only an impression: I don’t think we have much in the way of hard statistical evidence. The sociolinguists may know its percentage incidence in MLE (see table below), but there’s not a lot of information available about other varieties.
I don’t think I ever say ðə ˈʔæpl̩ and so on myself. But I could be wrong.
At the age of 18, as I was picking up German by staying with a family in northern Germany on a family exchange, I noticed that when wanting to know the time my exchange partner, rather than ask Wie viel Uhr ist es? (‘how many o’clock is it?’), as shown in my tourist’s phrasebook, would usually go for the formula Wie spät ist es? (‘how late is it?’). So I did so too.
Imitating his pronunciation, I pronounced spät as ʃpeːt, using the same vowel sound as in Wie geht’s viː ˈɡeːts ‘how’s it going?’.
As I got to grips with the written as well as the spoken language, I learnt to treat the umlauted letter ä as being pronounced exactly the same as the letter e.
Years later, when I studied phonetics with John Trim at Cambridge, he told me that the German pronunciation I had acquired through total immersion, while commendably native-like in its way, was in some respects regional. If I wanted to speak proper Hochdeutsch, I ought to remember to say Guten Tag! with taːk, not ta(ː)x; the train, der Zug, should be tsuːk, not tsʊx; and for long ä, as in spät, I ought to add a new item to my German vowel system, namely the long ɛː, thus ʃpɛːt.
The standard set out in German dictionaries and textbooks treats orthographic e and ä as having the same value when short, ɛ, but different values when long, namely eː and ɛː respectively.
So fällen ‘to fell’ ˈfɛlən is a perfect rhyme for bellen ‘to bark’ ˈbɛlən (both have the short vowel). But wählen ‘to choose’ should not, in Hochdeutsch, be a perfect rhyme for fehlen ‘to be lacking’ (with the the long vowel): ˈvɛːlən, ˈfeːlən.
This distinction still feels artificial to me, and I don’t make it unless perhaps carefully reading some text aloud or making a phonetic point.
The pronunciation dictionaries tend to hedge their bets on this distinction. Here’s the sixth edition of the Duden Aussprachewörterbuch. Der Vokal [ɛː] kann auch [eː] gesprochen werden… (p. 21: ‘The vowel [ɛː] can also be pronounced [eː]…’)
And here’s the Deutsches Aussprachewörterbuch. Der Unterschied zwischen [eː] und [ɛː] wird in der Aussprache meist nich stark verdeutlicht, so dass häufig ein Vokalklang zwischen [eː] und [ɛː] mit einer Tendenz zu [eː] entsteht. (p. 58: ‘The difference between [eː] and [ɛː] is for the most part not made very clearly in pronunciation, so that frequently a vowel quality between [eː] and [ɛː] arises, with a tendency towards [eː].’)
Wikipedia says, I think quite correctly, The long open-mid front unrounded vowel [ɛː] is merged with the close-mid front unrounded vowel [eː] in many varieties of Standard German…
I shall continue to speak German with an undifferentiated eː.

One or two of the people commenting on nt-reduction (blog, 18 Nov.) also mentioned the possibility of twenty having the vowel ʌ.
Kensuke Nanjo said According to my daily observation of American English, I think this variant is worth including in pronouncing dictionaries. Quite a few Americans use it and as you may know, this is the second variant for "twenty" in the Merriam-Webster Collegiate Dictionary.
There are indeed plenty (“plunny”?) of Americans who seem to pronounce twenty with a seriously backed and lowered quality as compared with their default DRESS vowel.
However, in deciding whether this is a sporadic irregularity found just in this word (and perhaps in plenty too), we must first establish what is their default DRESS vowel. We need to discount the possible effects of what, following Labov, has come to be known as the Northern Cities Vowel Shift.
The “northern cities” (of America) in which this sound change flourishes are clustered around the Great Lakes: places such as Buffalo, Cleveland, Detroit, and Chicago. The geographical extent of the shift varies depending on which vowel is involved and in which phonetic environment(s); and in any case it is also socially and stylistically variable. But what it can do is to make DRESS words sound as if they have the STRUT vowel — perhaps all of them, perhaps particularly those in which the vowel is followed by a nasal. Note that the STRUT vowel shifts too, so that we do not normally get loss of the distinctions exemplified in get – gut, bed – bud, wren – run etc.
So someone who says ˈtwɛ̈ni, with a thoroughly retracted vowel, is not necessarily saying ˈtwʌni (“twunny”), to rhyme with funny.
Others, though, are. They include rirelan, who mentioned twenty: /ˈtwʌni/ (along with "plenty" /ˈplʌni/. "plentiful" is still /ˈplɛntəfəl/ though.)
Furthermore, Americans from other, mainly southern or western, parts of the country may merge pen and pin as pɪn (i.e. merge DRESS with KIT before a nasal). For them, twenty may rhyme, if not with funny, then with skinny as well as with many.
Kensuke reckons that a reasonably exhaustive pronunciation dictionary ought to give AmE twenty as ˈtwenti, ˈtwʌnti, ˈtweni, ˈtwʌni. Seems reasonable, though perhaps we ought to add ˈtwɪnti, ˈtwɪni, too.