Micro Speaker for Voice-Enabled Consumer Electronics

Writer:By Shenzhen Hongsheng Electronic Industry Co. LTD Visits: 10 09, 2026

Micro Speaker for Voice-Enabled Consumer Electronics

A consumer product that listens to the user has a speaker problem that a music product does not. The microphone decides when the device responds, but the speaker decides whether the response is understood, whether the user believes it was aimed at them rather than at the room, and whether they speak again. Voice interaction replaces a frequency response requirement with a comprehension requirement, and that changes which dimensions of a micro speaker matter first.

A consumer product that listens to the user has a speaker problem that a music product does not. The microphone decides when the device responds, but the speaker decides whether the response is understood, whether the user believes it was aimed at them rather than at the room, and whether they speak again. Voice interaction replaces a frequency response requirement with a comprehension requirement, and that changes which dimensions of a micro speaker matter first.

A voice-enabled product is judged on a loop, not on a sound. The device hears an utterance, processes it, produces a spoken confirmation, and the user decides whether to continue. Every element of that loop can be measured, but only two of them are usually specified in the acoustics section of the requirement.

  1. What Changes When the Product Listens

Short answer:  A voice-enabled product is judged on a response loop. The speaker decides whether the confirmation is understood, whether it appears aimed at the user, and whether they speak again — not whether it sounds musical.

Three things go wrong in practice. The confirmation is intelligible on the bench at 10 cm and unintelligible across a room. The user speaks at 60 dB and the device answers at a level that is correct on the datasheet but masked by a fan, a motor or background speech. And the product ships with a microphone beamforming layout that assumes the speaker will not radiate broadly into the same space, which is a different product from the one where the speaker was chosen.

None of these three failures is visible in a sensitivity figure. They appear at the listening position, in the real noise, through the real front face of the product.

  2. Why Voice Comprehension Is Not a Frequency Response Requirement

Short answer:  Intelligibility depends on keeping the critical speech band at a level high enough to clear the local noise floor. Low-frequency output adds voice body and naturalness, not comprehension, so two products with very different bass capability can be equally usable.

Intelligibility depends mainly on preserving the critical speech band at an adequate level, and on the ratio of that level to the noise at the listening position. Everything below the speech band contributes voice body and naturalness rather than comprehension. This is why two products with very different low-frequency capability can be equally usable for voice, and why a product that sounds thin can still be perfectly clear.

This also means that the published resonance of a driver is a weak predictor of voice performance. A product asking for a low resonance is often asking for something the application does not use, while a product with no low-frequency requirement at all may need a higher level in the speech band than its datasheet suggests.

Microphone array behaviour should be treated as an influence on the acoustic path, not as a specification of the driver. Speaker radiation pattern, enclosure, placement and mechanical isolation all affect the echo path that the array sees. There is no single driver property that makes an array work better; the geometry of the product decides that.

  3. How to State the Requirement So It Can Be Answered

Short answer:  Five items make a voice-speaker requirement answerable: the space for the driver and its cavity, the impedance, the available power, the target level at the real listening position, and the local noise floor. Without all five, any comparison is partial.

The target level has to be stated at the position the user will be at, with the distance and the noise floor behind it. A sensitivity quoted at 10 cm tells you nothing about a device used across a living room. Speech intelligibility measurement methods, such as those published in the ITU-T P-series, apply to the whole voice path and cannot be reduced to a single driver figure.

1. State the target level at the position the user will be at, with the distance and the noise floor behind it, rather than at a bench distance.

2. State the frequency band the response has to cover for the intended content — voice confirmation, or voice plus music — because the two lead to different drivers.

3. Ask what cavity volume is actually available behind the driver in the assembled product, and at what measurement volume the supplier quoted the sensitivity.

4. Ask what the front face of the product does acoustically, and where the microphone sits relative to the speaker and to the user.

5. Ask for the configuration to be evaluated in the product's own cavity, not only in a test box, because the two can differ substantially.

6. Ask which part of the acoustic design would have to change first if the target cannot be met with the specified driver — that answer usually identifies the real constraint.

  4. Driver Classes Used in Voice-Enabled Products

Short answer:  Published catalogue data separates into three useful classes: plain round drivers with a rear cavity, drivers needing a defined host cavity (BOX construction), and modules that bring their own volume. A fourth group trades height for level and suits products where the front face is already occupied.

Table 1: Illustrative class comparison observed across published micro speaker ranges. Values are catalogue figures at the stated test conditions, not a product recommendation.

Class

Acoustic path

Typical published sensitivity

Typical published F0

Where it fits a voice product

Round driver with rear cavity

Needs defined volume behind the diaphragm

93–98 dB at 2 kHz / 10 cm

500–1000 Hz ±15%

Products with a designed acoustic chamber and space behind the driver

Driver requiring a host cavity

BOX construction; the product front face becomes part of the acoustic path

95–99 dB at 2 kHz / 10 cm in the stated box

600–1050 Hz ±15%

Thin products where no rear volume can be reserved

Module with its own volume

Front-face output; cavity is part of the module

103–105 dB at 2 kHz / 10 cm

500–800 Hz ±15%

Pockets too small for a cavity, or products needing high level without rear space

High-height driver

Larger radiating structure; height traded against level

95–97 dB at 2 kHz / 10 cm

350–600 Hz ±15%

Products with generous internal height and a low target resonance

  5. Published Data for Voice-Oriented Configurations

Short answer:  The configurations below are taken from one published sample catalogue. Published figures are specified with their test conditions; each should be compared with the intended cavity and measurement distance before being treated as a prediction.

Table 2: Published catalogue configurations relevant to voice-enabled products, with stated test conditions.

Model

Format and published size

Published sensitivity

Published F0

Impedance

Stated application

HS001534H39

Round magnetic, φ15 × 3.4 mm

90 dB at 2 kHz / 10 cm / 0.8 W

800 Hz ±15%

8 ±15% Ohm

Voice remote controllers, smart home, door locks

HS002038H

Round magnetic, φ20 × 3.8 mm

93 dB at 2 kHz / 10 cm / 1.0 W

800 Hz ±15%

8 ±15% Ohm

POS, security and alarms, smart home

HS002045H

Round magnetic, φ20 × 4.5 mm

94 dB at 2 kHz / 10 cm / 1.0 W

600 Hz ±15%

8 ±15% Ohm

Smart devices with a defined sound chamber

HS0023123H123

Pot-type, φ23 × 12.3 mm, composite diaphragm, PU edge

95 dB at 2 kHz / 10 cm / 2.0 W

400 Hz ±15%

4 ±15% Ohm

Smart home, Bluetooth speakers and lamps

HS002850H50

φ28 × 5.0 mm iron frame

97 dB at 2 kHz / 10 cm / 2.0 W

600 Hz ±15%

8 ±15% Ohm

Smart home, security alarms, learning machines

HS0028110H110

Pot-type, φ28 × 10.0 mm

96 dB at 2 kHz / 10 cm / 2.0 W

350 Hz ±15%

4 ±15% Ohm

Smart home, learning machines, Bluetooth speakers

HS003021H

φ30 module, 21 mm height

105 dB at 2 kHz / 10 cm / 2.0 W

800 Hz ±15%

4 ±15% Ohm

AI voice products, voice intercom

HS003058H

φ30 module, φ15.5 mm core

103 dB at 2 kHz / 10 cm / 2.0 W

800 Hz ±15%

4 ±15% Ohm

Feature phones, voice and smart home

HS-BX-283115H

28 × 31 × 15 mm, φ15.5 mm core

97 dB at 2 kHz / 10 cm / 2.0 W

640 Hz ±15%

4 ±15% Ohm

AI voice products

HS-BX-282813H

28 × 28 × 13 mm, φ15.5 mm core

97 dB at 2 kHz / 10 cm / 2.0 W

880 Hz ±15%

4 ±15% Ohm

AI voice products

HS-BX-203008H

20 × 30 × 8 mm, φ12.5 mm core

96 dB at 2 kHz / 10 cm / 3.0 W

1000 Hz ±15%

4 ±15% Ohm

All-in-one and industrial control machines

HS201623H

20 × 16 × 2.3 mm, no front cover, BOX use only

96 dB at 2 kHz / 10 cm / 1.0 W / 1 cc box

800 Hz ±15%

7 ±15% Ohm

Smartphones and similar audio

HS201623H is listed with a no-front-cover design for BOX use only; it requires a defined host cavity and cannot be fitted as a bare driver against an open front face.

Project Case Study (Hongsheng)

A voice-control unit for a home entrance was specified with three targets at once: a resonance at or below 300 Hz, a sensitivity 4 dB above what the customer had budgeted, and two drivers inside a 16.5 mm thickness limit. The first difficulty was structural rather than electrical: two drivers in that envelope needed one shared acoustic cavity, so Hongsheng designed an integrated pair sharing a single cavity and added a compliant surround height of 1.5 mm to allow side output within the same thickness. The measured result was a published sensitivity of 81 ± 3 dB at 2 kHz, 2.83 Vrms input, 1 m, and a resonance of 420 Hz ±15% at a rated power of 5.0 W and a maximum of 8 W, with the complete two-driver assembly and its cavity supplied at RMB 12.5 per set. Side output limited the effective bandwidth to 3 kHz, which the customer accepted because the product is a voice broadcast unit; the target resonance was not reached, and the customer accepted 420 Hz on the same reasoning that the written specification placed no requirement below 300 Hz. Hongsheng can supply the integrated cavity pair, or a single-driver alternative where the acoustic space is larger, once the available envelope and the target level at the listening position are stated.

One shared cavity carried two drivers inside 16.5 mm, and the written specification — not the datasheet — decided which targets had to give way.

Hongsheng's catalogue covers all four of these classes, which matters for a voice product because the same intelligibility requirement can be met by different classes depending on how much internal depth the product can reserve. Hongsheng can evaluate the classes against a stated envelope and target level, and can state plainly where the requirement cannot be met as written.

  6. What to Confirm Before Ordering

Short answer:  Confirm the level at the real listening position, the band the content needs, the cavity behind the driver, the front-face construction, and which parameter would have to change first. These five answers determine the class, and most disputes come from unanswered ones.

A voice requirement is unusually easy to state accurately and unusually easy to state inaccurately. Once the five items above are written down with numbers, the remaining work is ordinary acoustic design.

  7. Frequently Asked Questions

Q1: Does a voice-enabled product need a micro speaker with a low resonance?

Not necessarily. Comprehension depends mainly on the critical speech band and on the level relative to the local noise floor. A higher resonance is often acceptable when the product is a confirmation-voice device; a lower resonance becomes relevant when the same speaker also carries music or alerts.

Q2: What target level should I state for a voice confirmation?

State it at the position the user will be at, together with the distance and the noise floor at that position. A sensitivity quoted at 10 cm in a test box does not describe that condition, and quoting it as if it does is the most common cause of a speaker that is correct on the bench and unusable in the product.

Q3: Does a microphone array require a driver with a narrow radiation pattern?

Treat it as an influence rather than a requirement. Speaker radiation pattern, enclosure, placement and mechanical isolation all affect the acoustic echo path seen by the array. No single driver property makes an array behave better; the geometry of the product decides that.

Q4: Is a BOX-type module easier to specify than a driver with a rear cavity?

It removes one unknown, because the cavity is part of the module and the published figure already refers to it. The trade-off is that the front face of the product becomes part of the acoustic path, so the module's published performance still depends on how the front face is designed.

Q5: How should I compare two published sensitivity figures?

Check the measurement distance, the applied voltage and the cavity volume behind the driver. Without all three, the figures are not comparable. Two datasheets can differ by several decibels on paper while describing different conditions rather than different products.

Q6: Which standard should be used to measure speech intelligibility?

Methods published in the ITU-T P-series for speech intelligibility apply, such as ITU-T P.807. Recommendations on subjective or objective speech quality are not interchangeable with intelligibility measurement and should not be cited as if they were.

Q7: What if the target level and the target resonance cannot both be met?

Ask the supplier to state which part of the acoustic design would have to change first. That answer identifies the real constraint — cavity volume, front-face area, height or power rail — and turns an impossible requirement into a decision.

  8. Summary

A voice-enabled product changes what a micro speaker has to achieve, without changing what a micro speaker is. Comprehension is a function of the speech band and of the level relative to local noise; low-frequency capability adds naturalness rather than intelligibility. Selecting by published resonance alone therefore tends to over-specify the wrong parameter. Hongsheng can evaluate voice-speaker requirements against the available envelope, the target level at the listening position and the intended cavity, and state plainly which parameter would need to move if the first proposal does not reach the target. Final selection always belongs in the product specification and in measurements made in the product itself.

Next step  If you are evaluating a micro speaker for a voice-enabled consumer product, the shortest route to a configuration worth testing is to state five things: the space available for the driver and its cavity, the impedance the amplifier will drive, the power the rail can supply, the target level at the position the user will actually listen from, and the noise floor at that same position. With those specified, our engineering team can recommend a suitable configuration for evaluation, or state plainly which part of the acoustic design has to change first.

More in This Series

· How a shallow housing changes the result, and when a BOX platform is the right route → https://www.hsdz-spk.com/news/575.html

· Voice-enabled products are judged on comprehension, not on a frequency curve → https://www.hsdz-spk.com/news/576.html