AI Robot Speaker Design: Acoustic Echo, Microphone Proximity and Self-Voice
Published: 2026-10-05 | Use case: the second acoustic direction in a voice product — keeping the microphone able to hear the user while the product is speaking
The question a voice-product team actually asks is not whether the speaker can play loud enough. It is whether the robot can speak while it still hears the user. Those two requirements pull against each other through the enclosure, because the microphone and the driver share whatever volume the product gives them, and every decibel of self-voice the microphone has to reject is a decibel of user speech it is working against. The symptoms show up in the software — a product that answers itself, a wake word that fires when nobody spoke, a user who is interrupted mid-sentence — and are usually investigated in the signal chain, where the cause is rarely to be found. It is an acoustic problem, and it is decided by proximity, enclosure and drive condition before it is decided by firmware.
1. The Product Has Two Acoustic Jobs
Short answer: A voice product must project to the user and still leave the microphone able to hear the user, and the second job depends on enclosure rather than on output.
The obvious requirement is level: the assistant has to be heard across a room, from a kitchen, or from the other side of a desk. The less obvious requirement is that the microphone must still work while the assistant is speaking, and must not treat the product's own output as a user command. These two requirements are not independent, and they interact in two places. They interact through the enclosure, because the microphone and the driver share whatever volume the industrial design gave them. They interact through the signal chain, because cancellation has a finite budget and everything the microphone picks up from the product spends part of it. A design that specifies the driver first and the microphone placement later has already chosen which requirement to sacrifice, usually without noticing that the decision was made.
Table 1: The two directions of a voice product and what each one needs
Direction | What it needs from the driver and enclosure | What goes wrong when it is neglected |
Product to user | Level at the listening distance, speech-band clarity | The prompt is inaudible, so the user reaches for a button instead |
User to microphone | Low self-radiation into the enclosure, low distortion | Wake-up fails, or the product talks over the user |
Both simultaneously | Separation, controlled radiation, and a cavity the microphone does not see | The product answers itself, or ignores a user speaking over it |
2. The Speech Band Is the Acceptance Criterion
Short answer: A voice prompt should be accepted on intelligibility in the 2–4 kHz region at a moderate level, not on a maximum sensitivity figure.
A published maximum is easy to quote and easy to misread. What a voice product needs is a level and a spectral balance that keep the speech band clear once the enclosure and the microphone position have been imposed. For prompts, the working target is a 2–4 kHz region reproduced at a moderate sound pressure level, because that is where the consonant information that carries intelligibility lives. Two consequences follow. A driver whose useful output sits higher is working against the product, because the region that matters is the region a small shallow enclosure and a front mesh attenuate first. And an intelligibility assessment is a better acceptance criterion than a sensitivity figure, because it measures the outcome the user experiences rather than a property of the part. Published methods for that assessment exist in the ITU-T P-series, and ITU-T P.807 is the applicable recommendation for speech intelligibility measurement.
3. Microphone Proximity and the Echo Path
Short answer: Acoustic echo cancellation assumes a reasonably predictable path from the driver to the microphone, and proximity makes that path less predictable rather than simpler.
On a desk product the microphone and the driver are often centimetres apart, separated by the product's own walls, and the acoustic path between them changes with wherever the user places the product. The closer they are, the more the microphone hears and the more the cancellation has to work, and the more the result depends on the acoustic detail of the enclosure. This is why the microphone position is decided with the driver rather than after it. It is also why the component property should be described carefully: speaker radiation pattern, enclosure, placement and mechanical isolation all affect the acoustic echo path seen by the microphone array, and none of them on its own determines whether an array can work. Treating driver directivity as a precondition for beamforming overstates the case; it is one factor among several in a path that is mostly determined by the product's geometry.
4. Self-Voice and the Self-Triggering Wake Word
Short answer: When the product's own playback reaches the microphone, a wake phrase inside the prompt can be detected as a command, and the failure looks like a software bug.
The classic symptom is a product that answers itself: it speaks a confirmation containing a phrase the wake-word engine recognises, the engine fires, and the product responds again. Engineers typically investigate the wake-word threshold first, and the threshold is rarely the cause. The cause is usually that the self-voice level at the microphone is close enough to the user-voice level that no comfortable threshold separates them. There are three ways out and they are complementary rather than alternative: reduce the playback reaching the microphone through separation, orientation and radiation control; shape the confirmation so it is not recognisable as a command where the signal chain allows it; and use the fact that the product knows when it is speaking. The acoustic part is the first of the three and it has to exist before the firmware does, because everything downstream is compensating for a physical arrangement.
5. Voice Body Is Not the Same Thing as Intelligibility
Short answer: Low-frequency output contributes to voice body and naturalness, but intelligibility depends much more on preserving the critical speech-band information.
The two are routinely conflated, and conflating them produces the wrong specification. A product that needs prompts to sound natural and have apparent quality benefits from low-frequency content, because that is what gives a voice its body and warmth, and it is why a ported arrangement is worth considering where the enclosure can be vented safely. What that same content does not do is carry intelligibility. Intelligibility is carried by the speech band, and the correct engineering response is to protect that band through the enclosure, the grille and the drive condition rather than to chase low-frequency extension and assume the product will follow. A sealed cavity stiffens the air spring and raises the system resonance according to FC = Fs × √(1 + Vas / Vb), so a product that wants both voice body and a low resonance has to reach for a ported or boxed route rather than expecting a thin sealed part to deliver both.
6. Three Enclosure Routes
Short answer: Voice products resolve into a thin driver, a standard box or a high-output box, and the choice follows from listening distance, available volume and microphone separation.
Table 2: Candidate parts by route and source character (values as published in the supplier's sample catalogue, subject to the product datasheet)
Model | Size | SPL (stated condition) | F0 | Route and source character |
HS003021H | φ30 BOX, 21 mm height | 105 dB @ 2 kHz / 10 cm / 2.0 W | 800 Hz | High-output BOX; cavity defines the radiation |
HS003058H | φ30 BOX | 103 dB @ 2 kHz / 10 cm / 2.0 W | 800 Hz | High-output BOX; large magnet design |
HS-BX-0045-KT5 | φ40 BOX, 14 mm height | 103 dB @ 2 kHz / 10 cm / 2.0 W | 500 Hz | High-output BOX; lower resonance for voice body |
HS002628H28 | φ26 BOX, 28 mm height | 99 dB @ 2 kHz / 10 cm / 2.0 W | 500 Hz | Standard BOX; deep cavity |
HS-BX-283115H | 28×31×15 mm BOX | 97 dB @ 2 kHz / 10 cm / 2.0 W | 640 Hz | Standard BOX; rectangular, cavity carried with the part |
HS-BX-282813H | 28×28×13 mm BOX | 97 dB @ 2 kHz / 10 cm / 2.0 W | 880 Hz | Standard BOX; same level, 240 Hz higher resonance |
HS-BX-1217-X10 | 1217 BOX, front sound output | See the product datasheet | See datasheet | Compact BOX platform for all-in-one devices |
HS-BX-203008H | 2030 BOX | 96 dB @ 2 kHz / 10 cm / 3.0 W | 1000 Hz | Standard BOX; higher drive, higher resonance |
HS002850H50 | φ28×5.0 mm | 97 dB @ 2 kHz / 10 cm / 2.0 W | 600 Hz | Bare round; radiation set by the host enclosure |
HS003650H | φ36×5.0 mm | 97 dB @ 2 kHz / 10 cm / 2.0 W | 500 Hz | Bare round; more low end, same pattern question |
HS005717H | φ57×17 mm | 108 dB @ 2 kHz / 10 cm / 2.0 W | 550 Hz | Bare round, high output; needs a larger body |
HS251233H | 25×12×3.3 mm | 97 dB @ 2 kHz / 10 cm / 2.0 W / 1 cc | 750 Hz | Thin bare driver; for flat bodies and close listening |
Table 3: Choosing between the three routes
Route | Suits | Gain | Cost |
Thin bare driver | Close listening, flat body, no cavity available | Lowest profile and cost | Resonance sits high; radiation set by the host |
Standard BOX platform | Desk and appliance products with a modest volume | Tuned cavity, reproducible radiation | Fixed footprint and depth |
High-output BOX | Room-filling voice, far-field, noisy environments | Level and voice body at distance | Volume, power and heat in the product |
Project Case Study (Hongsheng)
A projector programme needed television-class low-frequency performance from a thin, light body that the customer benchmarked against a completely different product class. The solution was not a larger driver: it was several rounds of acoustic tuning combined with a purpose-built cavity and a matched loudspeaker, so that the thin body produced the required output rather than being replaced by a thicker one. The result was acoustic performance equivalent to the benchmark, with 200,000 units shipped globally in batches, and the approach became a reference solution for other manufacturers building the same category. The transferable lesson for a voice product is that the enclosure route is a design variable, and the tuning rounds are what convert it into a result. Hongsheng's BOX speaker platforms are suited to applications where the cavity needs to be controlled as part of the speaker assembly, which is also the route that makes the acoustic behaviour reproducible between the bench and the product.
One-line conclusion: the thin body was made to behave like a larger one acoustically, and that took tuning rounds rather than a bigger part.
7. What to Confirm About the Echo Path and the Enclosure Route
Short answer: Seven items turn a voice speaker from a loudness figure into a two-way acoustic design, and most of them are about the enclosure rather than the part.
1. The 2–4 kHz output at the product's actual drive level, not the free-field maximum.
2. The measurement basis: frequency, distance, drive level and enclosure reference.
3. The radiation characteristic of the assembly, measured in the finished enclosure.
4. What the driver contributes to the sound field inside the enclosure, for microphone placement purposes.
5. The enclosure route options available, and the footprint and depth each one fixes.
6. The tolerance band on sensitivity and resonance, and the conditions it was measured under.
7. Whether an intelligibility assessment in the assembled product can be supported, and by what method.
Hongsheng can state the radiation and enclosure options available for a given voice product, and can supply the measured report for the shipped units so the echo path is documented rather than inferred.
8. FAQ — AI Robot Speaker Design
Q1: What is different about a voice product compared with a device that only plays prompts?
A: It has to satisfy two directions at once. The product must be heard at the listening distance with speech-band clarity, and the microphone must still hear the user while the product is speaking. The second requirement depends on proximity, enclosure and drive condition, and it is the one that is usually left unspecified.
Q2: How do I stop a voice product from waking itself up?
A: Three complementary moves — reduce the playback reaching the microphone through separation, orientation and radiation control; shape the confirmation so it is not recognisable as a command where the signal chain allows it; and use the fact that the product knows when it is speaking. The acoustic fix has to exist before the firmware does.
Q3: Does the driver need to be directional for a microphone array to work?
A: No. Speaker radiation pattern, enclosure, placement and mechanical isolation all affect the acoustic echo path seen by the array, and none of them alone determines whether an array can work. Treating directivity as a precondition overstates the case; it is one factor in a path mostly determined by the product's geometry.
Q4: Does low-frequency output improve speech intelligibility?
A: Not directly. Low-frequency output contributes to voice body and naturalness, but intelligibility depends much more on preserving the critical speech-band information. Chasing low-frequency extension in a sealed cavity will not improve the speech band, and in a small sealed cavity it raises the system resonance.
Q5: How should a voice product be accepted?
A: By an intelligibility assessment in the assembled product rather than by a sensitivity figure, because the enclosure, grille and microphone position have by then consumed much of the budget. ITU-T P.807 is the applicable published recommendation for measuring speech intelligibility.
More in This Series — Micro Speaker Acoustic Requirements by Device Class
This article is part three of a three-part technical series on micro speaker acoustic requirements for specific device classes. Check out the other two articles from this guide:
· Part 1 — Micro Speaker Reliability for Smart Locks: Temperature, Adhesive, Magnet and Cavity Design →https://www.hsdz-spk.com/news/565.html
· Part 2 — Video Doorbell Speaker Design: Balancing Acoustic Output and Ingress Protection →https://www.hsdz-spk.com/news/566.html
9. Summary — Design the Second Direction
A voice product's loudspeaker has two jobs, and the second is the one that decides whether the product is usable. Being heard across a room is a level requirement and it is straightforward. Letting the microphone still hear the user while the product speaks depends on proximity, enclosure, radiation and drive condition, and none of those are decided by the part number. Size the driver against the listening distance, define the acceptance criterion on the 2–4 kHz band rather than on a maximum figure, keep low-frequency content in its proper place as voice body rather than as an intelligibility lever, and treat the product's own output as a design input rather than as an accident. Hongsheng publishes BOX platforms for exactly this reason — when the cavity is controlled as part of the speaker assembly, the acoustic behaviour stays the same between the bench and the product, and a voice assistant stops answering itself.