AI Robot Speaker Design: Acoustic Echo, Microphone Proximity and Self-Voice

Writer:By Shenzhen Hongsheng Electronic Industry Co. LTD Visits: 10 06, 2026

AI Robot Speaker Design: Acoustic Echo, Microphone Proximity and Self-Voice

Published: 2026-10-05  |  Use case: the second acoustic direction in a voice product — keeping the microphone able to hear the user while the product is speaking

The question a voice-product team actually asks is not whether the speaker can play loud enough. It is whether the robot can speak while it still hears the user. Those two requirements pull against each other through the enclosure, because the microphone and the driver share whatever volume the product gives them, and every decibel of self-voice the microphone has to reject is a decibel of user speech it is working against. The symptoms show up in the software — a product that answers itself, a wake word that fires when nobody spoke, a user who is interrupted mid-sentence — and are usually investigated in the signal chain, where the cause is rarely to be found. It is an acoustic problem, and it is decided by proximity, enclosure and drive condition before it is decided by firmware.

  1. The Product Has Two Acoustic Jobs

Short answer:  A voice product must project to the user and still leave the microphone able to hear the user, and the second job depends on enclosure rather than on output.

The obvious requirement is level: the assistant has to be heard across a room, from a kitchen, or from the other side of a desk. The less obvious requirement is that the microphone must still work while the assistant is speaking, and must not treat the product's own output as a user command. These two requirements are not independent, and they interact in two places. They interact through the enclosure, because the microphone and the driver share whatever volume the industrial design gave them. They interact through the signal chain, because cancellation has a finite budget and everything the microphone picks up from the product spends part of it. A design that specifies the driver first and the microphone placement later has already chosen which requirement to sacrifice, usually without noticing that the decision was made.

Table 1: The two directions of a voice product and what each one needs

Direction

What it needs from the driver and enclosure

What goes wrong when it is neglected

Product to user

Level at the listening distance, speech-band clarity

The prompt is inaudible, so the user reaches for a button instead

User to microphone

Low self-radiation into the enclosure, low distortion

Wake-up fails, or the product talks over the user

Both simultaneously

Separation, controlled radiation, and a cavity the microphone does not see

The product answers itself, or ignores a user speaking over it

  2. The Speech Band Is the Acceptance Criterion

Short answer:  A voice prompt should be accepted on intelligibility in the 2–4 kHz region at a moderate level, not on a maximum sensitivity figure.

A published maximum is easy to quote and easy to misread. What a voice product needs is a level and a spectral balance that keep the speech band clear once the enclosure and the microphone position have been imposed. For prompts, the working target is a 2–4 kHz region reproduced at a moderate sound pressure level, because that is where the consonant information that carries intelligibility lives. Two consequences follow. A driver whose useful output sits higher is working against the product, because the region that matters is the region a small shallow enclosure and a front mesh attenuate first. And an intelligibility assessment is a better acceptance criterion than a sensitivity figure, because it measures the outcome the user experiences rather than a property of the part. Published methods for that assessment exist in the ITU-T P-series, and ITU-T P.807 is the applicable recommendation for speech intelligibility measurement.

  3. Microphone Proximity and the Echo Path

Short answer:  Acoustic echo cancellation assumes a reasonably predictable path from the driver to the microphone, and proximity makes that path less predictable rather than simpler.

On a desk product the microphone and the driver are often centimetres apart, separated by the product's own walls, and the acoustic path between them changes with wherever the user places the product. The closer they are, the more the microphone hears and the more the cancellation has to work, and the more the result depends on the acoustic detail of the enclosure. This is why the microphone position is decided with the driver rather than after it. It is also why the component property should be described carefully: speaker radiation pattern, enclosure, placement and mechanical isolation all affect the acoustic echo path seen by the microphone array, and none of them on its own determines whether an array can work. Treating driver directivity as a precondition for beamforming overstates the case; it is one factor among several in a path that is mostly determined by the product's geometry.

  4. Self-Voice and the Self-Triggering Wake Word

Short answer:  When the product's own playback reaches the microphone, a wake phrase inside the prompt can be detected as a command, and the failure looks like a software bug.

The classic symptom is a product that answers itself: it speaks a confirmation containing a phrase the wake-word engine recognises, the engine fires, and the product responds again. Engineers typically investigate the wake-word threshold first, and the threshold is rarely the cause. The cause is usually that the self-voice level at the microphone is close enough to the user-voice level that no comfortable threshold separates them. There are three ways out and they are complementary rather than alternative: reduce the playback reaching the microphone through separation, orientation and radiation control; shape the confirmation so it is not recognisable as a command where the signal chain allows it; and use the fact that the product knows when it is speaking. The acoustic part is the first of the three and it has to exist before the firmware does, because everything downstream is compensating for a physical arrangement.

  5. Voice Body Is Not the Same Thing as Intelligibility

Short answer:  Low-frequency output contributes to voice body and naturalness, but intelligibility depends much more on preserving the critical speech-band information.

The two are routinely conflated, and conflating them produces the wrong specification. A product that needs prompts to sound natural and have apparent quality benefits from low-frequency content, because that is what gives a voice its body and warmth, and it is why a ported arrangement is worth considering where the enclosure can be vented safely. What that same content does not do is carry intelligibility. Intelligibility is carried by the speech band, and the correct engineering response is to protect that band through the enclosure, the grille and the drive condition rather than to chase low-frequency extension and assume the product will follow. A sealed cavity stiffens the air spring and raises the system resonance according to FC = Fs × √(1 + Vas / Vb), so a product that wants both voice body and a low resonance has to reach for a ported or boxed route rather than expecting a thin sealed part to deliver both.

  6. Three Enclosure Routes

Short answer:  Voice products resolve into a thin driver, a standard box or a high-output box, and the choice follows from listening distance, available volume and microphone separation.

Table 2: Candidate parts by route and source character (values as published in the supplier's sample catalogue, subject to the product datasheet)

Model

Size

SPL (stated condition)

F0

Route and source character

HS003021H

φ30 BOX, 21 mm height

105 dB @ 2 kHz / 10 cm / 2.0 W

800 Hz

High-output BOX; cavity defines the radiation

HS003058H

φ30 BOX

103 dB @ 2 kHz / 10 cm / 2.0 W

800 Hz

High-output BOX; large magnet design

HS-BX-0045-KT5

φ40 BOX, 14 mm height

103 dB @ 2 kHz / 10 cm / 2.0 W

500 Hz

High-output BOX; lower resonance for voice body

HS002628H28

φ26 BOX, 28 mm height

99 dB @ 2 kHz / 10 cm / 2.0 W

500 Hz

Standard BOX; deep cavity

HS-BX-283115H

28×31×15 mm BOX

97 dB @ 2 kHz / 10 cm / 2.0 W

640 Hz

Standard BOX; rectangular, cavity carried with the part

HS-BX-282813H

28×28×13 mm BOX

97 dB @ 2 kHz / 10 cm / 2.0 W

880 Hz

Standard BOX; same level, 240 Hz higher resonance

HS-BX-1217-X10

1217 BOX, front sound output

See the product datasheet

See datasheet

Compact BOX platform for all-in-one devices

HS-BX-203008H

2030 BOX

96 dB @ 2 kHz / 10 cm / 3.0 W

1000 Hz

Standard BOX; higher drive, higher resonance

HS002850H50

φ28×5.0 mm

97 dB @ 2 kHz / 10 cm / 2.0 W

600 Hz

Bare round; radiation set by the host enclosure

HS003650H

φ36×5.0 mm

97 dB @ 2 kHz / 10 cm / 2.0 W

500 Hz

Bare round; more low end, same pattern question

HS005717H

φ57×17 mm

108 dB @ 2 kHz / 10 cm / 2.0 W

550 Hz

Bare round, high output; needs a larger body

HS251233H

25×12×3.3 mm

97 dB @ 2 kHz / 10 cm / 2.0 W / 1 cc

750 Hz

Thin bare driver; for flat bodies and close listening

Table 3: Choosing between the three routes

Route

Suits

Gain

Cost

Thin bare driver

Close listening, flat body, no cavity available

Lowest profile and cost

Resonance sits high; radiation set by the host

Standard BOX platform

Desk and appliance products with a modest volume

Tuned cavity, reproducible radiation

Fixed footprint and depth

High-output BOX

Room-filling voice, far-field, noisy environments

Level and voice body at distance

Volume, power and heat in the product

Project Case Study (Hongsheng)

A projector programme needed television-class low-frequency performance from a thin, light body that the customer benchmarked against a completely different product class. The solution was not a larger driver: it was several rounds of acoustic tuning combined with a purpose-built cavity and a matched loudspeaker, so that the thin body produced the required output rather than being replaced by a thicker one. The result was acoustic performance equivalent to the benchmark, with 200,000 units shipped globally in batches, and the approach became a reference solution for other manufacturers building the same category. The transferable lesson for a voice product is that the enclosure route is a design variable, and the tuning rounds are what convert it into a result. Hongsheng's BOX speaker platforms are suited to applications where the cavity needs to be controlled as part of the speaker assembly, which is also the route that makes the acoustic behaviour reproducible between the bench and the product.

One-line conclusion: the thin body was made to behave like a larger one acoustically, and that took tuning rounds rather than a bigger part.

  7. What to Confirm About the Echo Path and the Enclosure Route

Short answer:  Seven items turn a voice speaker from a loudness figure into a two-way acoustic design, and most of them are about the enclosure rather than the part.

1. The 2–4 kHz output at the product's actual drive level, not the free-field maximum.

2. The measurement basis: frequency, distance, drive level and enclosure reference.

3. The radiation characteristic of the assembly, measured in the finished enclosure.

4. What the driver contributes to the sound field inside the enclosure, for microphone placement purposes.

5. The enclosure route options available, and the footprint and depth each one fixes.

6. The tolerance band on sensitivity and resonance, and the conditions it was measured under.

7. Whether an intelligibility assessment in the assembled product can be supported, and by what method.

Hongsheng can state the radiation and enclosure options available for a given voice product, and can supply the measured report for the shipped units so the echo path is documented rather than inferred.

  8. FAQ — AI Robot Speaker Design

Q1: What is different about a voice product compared with a device that only plays prompts?

A: It has to satisfy two directions at once. The product must be heard at the listening distance with speech-band clarity, and the microphone must still hear the user while the product is speaking. The second requirement depends on proximity, enclosure and drive condition, and it is the one that is usually left unspecified.

Q2: How do I stop a voice product from waking itself up?

A: Three complementary moves — reduce the playback reaching the microphone through separation, orientation and radiation control; shape the confirmation so it is not recognisable as a command where the signal chain allows it; and use the fact that the product knows when it is speaking. The acoustic fix has to exist before the firmware does.

Q3: Does the driver need to be directional for a microphone array to work?

A: No. Speaker radiation pattern, enclosure, placement and mechanical isolation all affect the acoustic echo path seen by the array, and none of them alone determines whether an array can work. Treating directivity as a precondition overstates the case; it is one factor in a path mostly determined by the product's geometry.

Q4: Does low-frequency output improve speech intelligibility?

A: Not directly. Low-frequency output contributes to voice body and naturalness, but intelligibility depends much more on preserving the critical speech-band information. Chasing low-frequency extension in a sealed cavity will not improve the speech band, and in a small sealed cavity it raises the system resonance.

Q5: How should a voice product be accepted?

A: By an intelligibility assessment in the assembled product rather than by a sensitivity figure, because the enclosure, grille and microphone position have by then consumed much of the budget. ITU-T P.807 is the applicable published recommendation for measuring speech intelligibility.

More in This Series — Micro Speaker Acoustic Requirements by Device Class

This article is part three of a three-part technical series on micro speaker acoustic requirements for specific device classes. Check out the other two articles from this guide:

· Part 1 — Micro Speaker Reliability for Smart Locks: Temperature, Adhesive, Magnet and Cavity Design →https://www.hsdz-spk.com/news/565.html

· Part 2 — Video Doorbell Speaker Design: Balancing Acoustic Output and Ingress Protection →https://www.hsdz-spk.com/news/566.html

  9. Summary — Design the Second Direction

A voice product's loudspeaker has two jobs, and the second is the one that decides whether the product is usable. Being heard across a room is a level requirement and it is straightforward. Letting the microphone still hear the user while the product speaks depends on proximity, enclosure, radiation and drive condition, and none of those are decided by the part number. Size the driver against the listening distance, define the acceptance criterion on the 2–4 kHz band rather than on a maximum figure, keep low-frequency content in its proper place as voice body rather than as an intelligibility lever, and treat the product's own output as a design input rather than as an accident. Hongsheng publishes BOX platforms for exactly this reason — when the cavity is controlled as part of the speaker assembly, the acoustic behaviour stays the same between the bench and the product, and a voice assistant stops answering itself.