What is a phoneme? Definition and the 44 English phonemes
- Written by
- Jack Limebear
- Published
- Last updated
ListenListen to this article
A phoneme is the smallest unit of sound in a language that can change the meaning of a word. When you swap a single phoneme, you can produce an entirely different word. /p/ and /b/ collocated with /æt/ transform to ‘pat’ and ‘bat’, two completely different words.
Specifically in English, phonemes are the reason that spelling and sound often disagree. A word like “fox” has four distinct phonemes that make it up (/f/ /ɒ/ /k/ /s/) but only three letters while ‘ship’ has four letters but only three phonemes (/ʃ/ /ɪ/ /p/). Understanding phonemes is essential for both language learning and understanding how modern text to speech systems turn written words into natural-sounding speech.
This article covers exactly what is a phoneme with examples, provides the full chart of 44 English phonemes, and covers how AI voice technology uses phoneme-level control to get pronunciation right.
Summary:
- A phoneme is the smallest unit of sound in a spoken language that can change a word’s meaning.
- Most linguists count 44 phonemes in English, made up of 24 consonant sounds and 20 vowel sounds.
- Phonemes describe sounds while graphemes are the letters or combination of letters that represent a phoneme in writing.
- The International Phonetic Alphabet (IPA) standardizes the symbols used to represent phonemes to make sure there is consistency across how these sounds are pronounced.
- Text to speech tools convert graphemes to phonemes before generating an audio file.
What is a phoneme?
A phoneme is the smallest unit of sound that can distinguish one word from another. Linguists identify phonemes based on finding two words that differ by exactly one sound. These words are called minimal pairs, with examples being:
- Pat / bat: /p/ and /b/ are the distinct phonemes in this minimal pair.
- Ship / sheep: /ɪ/ and /iː/ are the phonemes that change these words.
- Bat / bad: /t/ and /d/ are the distinct phonemes that mark the difference between these words.
If you find a word where changing one sound produces a new word, then those two sounds are separate phonemes in that language. If it produces a slightly different-sounding version of the same word, then they’re only variants of the same phoneme.
The word “phoneme” comes from the Greek phōnēma, which means utterance or sound made. This is a useful fact to remember when defining phonemes as they’re units of sound, not information about spelling. As you’ll see shortly, phonemes go beyond just being single letters and encompass entire aural sounds.
How many phonemes are in English?
In English, there are 44 phonemes across 24 consonant phonemes and 20 vowel phonemes. A slight distinction here is that this is only a general assumption. Depending on accent and how linguists choose to define different sounds, some believe there are as few as 42 or as many as 45. For example, a general American accent would pronounce the /r/ in “car”, while Received Pronunciation (British) doesn’t, which would create a different total count across these accents.
Other accents may merge vowels together, removing phonemes from the total count. Vowel mergers, such as how some Americans pronounce “cot” and “caught” the same, impact how linguists assess the total figures. Some marginal sounds, where diphthongs are seen as combinations rather than single phonemes, may also influence numbers. For example, some systems count /ʍ/ (the breathy sound at the beginning of 'which') as a distinct phoneme, while others treat it as a combination of /h/ and /w/.
The standard count is a 44-phoneme model. The exact breakdown within this figure is 24 consonants, 12 monophthong vowels, and 8 diphthongs, which is the standard used in UK phonetics teaching and in ESL instruction.
The 44 phonemes of English: full chart with examples
Consonant phonemes (24)
Vowel phonemes (12 monophthongs + 8 diphthongs = 20 total)
Symbols follow the British English (RP) system used in most phonics curricula; General American analyses differ slightly for the r-colored vowels.
What’s the difference between a phoneme and a grapheme?
A phoneme is a sound, while a grapheme is the written representation of that sound. A grapheme might be one letter (c-a-t), two letters (a digraph, like sh or ea), three (igh in “night”), or four (ough in “though”).
The mismatch between phonemes and graphemes is what makes spelling in English notoriously difficult. For example:
- One phoneme, many graphemes: The /iː/ sound can be spelled ee (sheep), ea (leaf), e (be), ey (key), ie (chief), or ei (ceiling).
- One grapheme, many phonemes: The letters ough represent different phonemes in "through" (/uː/), "though" (/əʊ/), "tough" (/ʌf/), and "thought" (/ɔː/).
Although the English language has 44 phonemes, it has over 200 graphemes to spell them. If you think of a more phonetic language like Spanish, where mapping is nearly one-to-one, English becomes extremely complex to speak and transcribe. It’s also why a text to speech model can’t simply read the letters out, which we’ll come back to shortly.

Phonemes vs. allophones
Allophones are variant pronunciations of the same phoneme that don’t change the word’s meaning. If there isn’t a minimal pair where a phoneme creates two different words, then you instead have an allophone.
For example, the /p/ in “pat” sounds different from the unaspirated version in “spat”. These are different sounds, but no English word pair is distinguished by this aspiration alone, meaning English speakers hear them as essentially the same sound. In different languages, like Thai, the difference between aspirated and unaspirated /p/ create separate phonemes that distinguish words completely.
What that means is that what counts as a phoneme is completely language-specific. It’s about what sounds do to the meaning of words, rather than solely acoustics.
How phonemes power text to speech
Every text to speech system needs to solve the same problem that a brain does when reading aloud: converting graphemes into phonemes before speaking. Grapheme-to-phoneme (G2P) conversion is notoriously difficult and is where many rudimentary systems make mistakes with TTS.
Text can be ambiguous. The exact same word, “read” could be /riːd/ in the present tense and /red/ in the past. Some proper nouns, like brand names or technical terminology, might not follow standard spelling rules at all. Modern AI models like Eleven v3 use context to resolve most of these, but you can also control phonemes more granularly for guaranteed pronunciation:
- Pronunciation dictionaries: Define how every word should be spoken across an entire project. Ideal for brand names, character names, medical terms, or company-specific vocabulary that appear repeatedly.
- SSML phoneme tags and IPA: You can get word-level control over pronunciation with SSML phoneme tags on v2 models, or by writing words directly with Eleven v3 audio tags.
The same phonemic layer runs in the opposite direction, too. For speech to text, models need to learn acoustic patterns so they can map the phoneme sequences before assembling them into words. Phonemic understanding is how they transcribe a word they may have never seen before.

How to use phonemes for language learning
One of the main reasons that a general user might want to know more about phonemes is because they’re extremely useful for language learning. A child’s ability to hear or identify individual phonemes at an early age directly influences their acquisition of further phonological knowledge and linguistic production throughout their adolescence.
When learning a new language, your brain has to compute phonemes that your ear was never trained to hear. Here are a few practical strategies to use phonemes to improve language learning:
- Learn IPA symbols: The International Phonetic Alphabet appears within dictionaries to demonstrate how to pronounce each word. Learning IPA means you’ll always know how to read a word, even if you’re seeing it for the first time.
- Listen and imitate phonemes: Slow down speech to identify each phoneme within a language and mirror it. Doing so helps build up context of how you should pronounce each sound, which is more effective than repeating entire sentences when learning languages.
- Use minimal pairs for refinement: Minimal pairs only have a singular phoneme change between them. Drilling examples of these in your speech will help train your tongue and ear to recognize and reproduce even subtle changes between languages.
AI voice tools have also made language learning significantly easier. With a text to speech generator, you can hear any word or phrase in a native-sounding voice in dozens of languages and numerous regional accents.
Try ElevenLabs for phoneme-level text to speech
Whether you're delving into the phonetics of a new language or refining your native tongue, a solid understanding of phonemes is the key to mastering the art of effective and articulate expression. Yet an academic understanding of linguistics is not the only way we humans can acquire languages. We also learn through listening and mimicry.
Text to speech generation tools are invaluable in cultivating phonemic awareness and building an innate understanding of sound units, whole words, and the other different sounds that are naturally included in spoken communication.
Get started with powerful text to speech technology with ElevenLabs or discover more about our suite of creative tools today.

.webp&w=3840&q=80)
.webp&w=3840&q=80)

