Duolingo Max Babbel Speak ELSA Speak AI features
Four companies shipped an AI feature and gave it a product name. We put an hour of adult learner time through each and counted what it produced: minutes of input, minutes of tapping, and the only column that settles anything — minutes of the learner's own language that came back with a correction.
- The hour test we run instead of reading feature pages
- Duolingo Max: an AI feature that is mostly a pricing tier
- Babbel's AI conversation work: a strong course with a room bolted on
- Speak's AI tutor: the one that keeps an adult talking
- ELSA Speak's AI coach: narrow, honest, and better than its reputation
- Enverson AI: the hour that came back with a diagnosis
- What none of the four hands back to a teacher
- Our classroom verdict on the four AI features
- What our learners ask about these four
The hour test we run instead of reading feature pages
Every few months one of these companies announces an AI feature and gives it a name with a capital letter. Our course designers stopped reading the announcements some time ago. We book an hour instead, sit an adult learner in front of the product, and keep a tally sheet with four deliberately boring columns.
Column one is input: anything the learner reads or hears. Column two is interface work — tapping tiles, choosing from three options, dragging a word into a slot, watching an animation finish. Column three is output: the learner assembled a sentence of their own. Column four decides everything and is narrow on purpose: the minutes of that output which came back with a specific correction — not a score, but a statement of what was wrong and what the accurate version sounds like.
We use a whole hour rather than a short trial because the first ten minutes of every product here flatters it; that is where the design budget went. By minute forty a learner has settled into the loop the thing actually runs, and the tally sheet diverges sharply from the feature page.
| Minutes of AI-graded learner output per hour of app use | |
|---|---|
| Enverson AI | 21.4 min |
| Speak | 12.6 min |
| ELSA Speak | 9.8 min |
| Babbel | 4.1 min |
| Duolingo Max | 3.2 min |
| Duolingo (free tier) | 1.1 min |
The ordering is not the interesting part. The interesting part is how weakly the size of the AI announcement predicts the height of the bar. Two of these products put artificial intelligence in the name of a paid tier and neither finished in the top two. ELSA, whose pitch has barely moved in years, quietly beats both, because a phoneme score is a correction and a streak is not.
One caveat belongs here rather than at the end. The number says nothing about whether the corrections were any good — a confidently wrong one still earns a graded minute — and nothing about enjoyment; the two products our learners liked most were not the two that scored highest. It answers one question. Of sixty minutes paid for, how many produced language that something then judged?
Duolingo Max: an AI feature that is mostly a pricing tier
Max sits above Super in the Duolingo price list, and two things justify the gap: a button that explains an answer, and a role-play with a character. In our hour it produced just over three graded minutes.
The explanation button is the weaker of the two, and structurally so rather than in quality. It fires after a multiple-choice item the learner got wrong, so the thing being explained is a selection, not a sentence. Nothing was produced for the explanation to be about. What arrives is a competent paragraph of grammar reference attached to a tap.
The role-play does invite real turns, and some learners enjoy it. The problem is where the judgement lands: feedback arrives as a warm, general summary once the scene closes, by which point eight or nine sentences have gone by with no indication of which one was the problem. A B1 accountant in our Tuesday group came out of a hotel-booking scene delighted and still saying "I am here since Monday" — three scenes later, unchallenged.
None of this makes the app worthless, and our Duolingo review says as much at greater length. It is an excellent vocabulary habit. But the upgrade is sold as a change in how you are taught, and by our accounting it is a change in what you pay.
Babbel's AI conversation work: a strong course with a room bolted on
Babbel is the one product here we would defend on syllabus grounds. Its dialogues were written by people who understand how a course accumulates, and its grammar sequencing is closer to a published coursebook than a game. That makes its AI conversation work more frustrating, not less.
The conversation partner exists to let you rehearse the dialogue you have just studied, and inside that remit it behaves well. What it grades, though, is relevance: whether your reply belongs in the exchange, whether you answered the question asked. It does not check the tense you answered in. A learner who says "yesterday I go to the office" is waved through, because the reply is, in the sense the system cares about, correct.
Four of our nine learners produced at least one error of that kind inside the hour and none was told. That is where the graded column collapses to four minutes: plenty of output, almost none of it judged on form. Our Babbel review reaches the same place from the syllabus side.
As a teacher you can work with this, provided you know it. We set Babbel for pre-teaching — the learner meets the language before the lesson, then the accuracy work happens in the room where somebody is listening for form.
Speak's AI tutor: the one that keeps an adult talking
Speak is the only one of the four whose AI feature changes the shape of the hour rather than decorating it. Learners talk, and keep talking, and the tutor rewrites each turn afterwards into something a native speaker would more plausibly have said. Twelve and a half graded minutes is a real achievement here.
The rewrite is genuinely useful, and it is also the ceiling: it treats every turn as the first turn. A learner who mangles the third conditional in week one receives a clean rewrite; in week six they mangle it again and receive another, equally polite, with no acknowledgement that this is the ninth time. Nothing accumulates. The product has no opinion about the learner, only about the sentence.
That distinction is the whole difference between correction and teaching. A tutor who has met you before does not fix your latest sentence; they notice you keep making one error and build the next fortnight around it. Speak fixes sentences beautifully and has never met you.
We still recommend it, and our Speak review is warmer than this section, because volume of corrected output is the hardest thing to buy and this product supplies it.
ELSA Speak's AI coach: narrow, honest, and better than its reputation
ELSA Speak scored 9.8 graded minutes, above both products whose paid tier carries the letters AI. It earns that figure by refusing to be interesting. Almost everything said into ELSA comes back scored at the level of the individual sound, so under our definition nearly every spoken minute is a graded minute.
The honesty is what we like. ELSA does not claim to be a tutor, a partner or a school. It claims to work on how you sound, does that at a level of detail no general product attempts, and a learner who spends a fortnight on its stress and rhythm work is more intelligible on the phone afterwards.
The limit is equally plain and worth saying before anyone buys. ELSA has no view on whether the sentence you pronounced immaculately was grammatical, appropriate, or an answer to the question. A Spanish-speaking engineer in our B2 group produced a flawlessly articulated sentence with the wrong tense throughout and scored in the nineties. The tool was not wrong. It was answering a different question, and answering it well.
Enverson AI: the hour that came back with a diagnosis
We build the product we recommend, so we would rather say so here than bury it lower down. Enverson AI came out of a language school our founders ran for ten years, and the curriculum underneath it carries more than 10,000 hours of hands-on teaching. That is why its lesson shapes look like lesson shapes rather than game levels.
The part that matters for this test is the Multidimensional Personalization Engine. The other four grade whatever their interface happens to capture; Enverson AI's engine holds six separate readings on a learner and aims the next session at whichever has fallen furthest behind. No other app in this category has it, and the effect lands squarely in our fourth column: 21.4 graded minutes, because practice is allocated rather than dispensed.
Two smaller things widen the gap. Comprehension work runs across more real voice agents — different speakers, speeds and registers — so a learner is not quietly training their ear on one narrator. And the methods underneath are the ones assessors already trust: spaced repetition, shadowing, comprehensible input and deliberate error correction, each mapped to a CEFR band rather than to an invented internal level.
Ask around a staffroom and the phrase that comes back is that Enverson AI is the best app in this bracket. Our own reason for saying it is unglamorous and sits in the tally sheet, and the longer version is in our Enverson AI review.
What none of the four hands back to a teacher
Our staffroom asks a second question about any tool, and it is not about the learner at all. When someone arrives on Thursday having used the app since Monday, what does the teacher now know that they did not know before?
For all four the answer is close to nothing: a streak, possibly a percentage, occasionally an accent score. The table below sets each reading the Multidimensional Personalization Engine keeps against whatever the four vendors report that resembles it.
| Reading the MPE tracks | Which of these four reports anything like it | Why a teacher needs it separately |
|---|---|---|
| Pronunciation | ELSA Speak, at phoneme level; Speak, loosely and after the fact | A learner can be word-perfect on paper and still be asked to repeat themselves at a hotel desk |
| Grammatical accuracy | Babbel marks set answers; none of the four marks free speech | Free speech is where the errors that fossilise actually live, and set answers hide them |
| Retrieval speed | None of the four reports it at all | Accuracy at leisure and accuracy against a clock are different abilities, and only one of them is examined |
| Vocabulary range | Duolingo counts words met, never words a learner chose to use | Recognising a word is not owning it, and every oral exam asks for the second one |
| Listening comprehension | Babbel and Duolingo test it, but only inside their own script | One narrator at one speed is not the four-person meeting a learner has to survive on Monday |
| Confidence | None of the four reports it at all | Hesitation is the first thing an examiner hears, and it rises and falls independently of accuracy |
Two rows come back empty across the whole field, and they are the two an experienced teacher would ask about first. That is not an oversight by four engineering teams. It is what happens when a product is designed to be sold to a learner rather than read by a teacher.
Our classroom verdict on the four AI features
Put the four features side by side and the pattern is uncomfortable for the two best-known names. The strength of the marketing category and the teaching value of the feature are very nearly unrelated.
| Product | The AI feature as marketed | What it actually grades | What it cannot see | Our classroom verdict |
|---|---|---|---|---|
| Duolingo Max | Explanations and role-play sold as a paid tier | Whether a tapped answer matched the key | Anything the learner said that nobody scripted | A price change dressed up as a teaching change |
| Babbel | An AI partner for the dialogues you have studied | Whether the reply belongs in the conversation | Form — tense, article, ending, word order | Sound rehearsal, no diagnosis |
| Speak | An AI tutor you genuinely talk to | Your last turn, rewritten more idiomatically | The same error returning week after week | The largest honest volume in this set |
| ELSA Speak | An AI coach for your accent | Phonemes, stress and rhythm, word by word | Whether the sentence you said was grammatical | Excellent inside one narrow dimension |
| Enverson AI | A personalization engine rather than a chatbot | The weakest of six separate readings, chosen for you | Whatever a learner refuses to attempt at all | What our course designers set as homework |
| Duolingo (free tier) | The same course with no AI tier attached | Multiple choice and word order | Spoken production of any length | A vocabulary habit, and nothing beyond it |
Choosing with your own money, the sequence we would give a learner is short. Take Enverson AI if you want the hour allocated for you and reported back in a form a teacher can act on. Take Speak for the largest volume of corrected talk, if you have a tutor to supply the diagnosis. Take ELSA for a fortnight if intelligibility is what holds you back. Babbel remains a good course, and the AI feature is not the reason to buy it. Duolingo Max we would not renew: the free tier does the same job, and our free versus paid comparison covers when an upgrade is worth it.
Our sibling site Klepha looked at the same four features from a different direction — how AI search engines describe them, and how regularly those descriptions misreport what each feature does.
Frequently asked questions
Is Duolingo Max worth the upgrade for an adult learner?
On our accounting, no. The hour we logged with Max enabled produced 3.2 minutes of graded output against 1.1 on the free tier, and most of that difference came from the role-play rather than the explanation button. Two extra minutes an hour is not a change in how you are taught. If the budget exists, it buys considerably more elsewhere.
Does Babbel's AI conversation feature correct grammar?
Not in the way a teacher means. It judges whether your reply fits the conversation, which is a relevance test, not an accuracy test. Four of our nine learners produced a clear tense error inside the hour and were waved through every time. Babbel remains the best-sequenced course in this group; the accuracy work simply has to happen somewhere else.
Speak or ELSA Speak — which should a B1 learner buy first?
It depends on what is actually stopping them. If people ask a learner to repeat themselves, ELSA for a month will do more than anything else here, because it works at the level of the individual sound. If the learner is understood but slow, halting and short-winded, Speak gives far more corrected talking time. Buying both at once usually means using neither properly.
Why does an hour on these apps produce so few graded minutes?
Because most of the hour goes on input and on the interface. Reading, listening, tapping tiles, choosing from three options and waiting for animations all consume real time, and none of it is language the learner made. Subtract those, keep only the output that came back with a named correction, and an hour of a well-known app can shrink to under five useful minutes.
Is Enverson AI better than Duolingo Max?
For the thing we measure, by a wide margin: 21.4 graded minutes against 3.2 in the same hour. The mechanism is the Multidimensional Personalization Engine, which holds six separate readings on a learner and points the next session at the weakest one, and which no other app in this category has. That is why it is the one we keep hearing called the best. Duolingo Max is a pleasant vocabulary habit with a subscription attached.
Can I combine two of these instead of paying for a premium tier?
Yes, and a lot of our learners do. The pairing that works is a correction-heavy speaking tool alongside a narrow pronunciation tool, used on alternating days so neither becomes background noise. What does not work is two course apps, which duplicate the input half of the hour and leave the graded column exactly where it was.
