Contents
You know the pattern: high agreement scores, thin free-text answers, and in the next town hall the workforce asks about exactly the topics the survey rated as good. That is not coincidence, it is design.
Three causes of shallow answers
- Social desirability. Employees answer the way they believe is expected, especially when anonymity is not credible.
- Scales without narrative space. A 1-to-5 scale asks for sentiment, not for reasoning. Without space for it, no reasoning comes back.
- Missing trust in follow-up. If no visible reaction has followed the last three surveys, people stop investing depth.
How to recognize the problem
| Signal | What it likely means |
|---|---|
| Agreement above 75% on all items | Acquiescence bias or social desirability |
| Free texts shorter than 10 words | Lack of trust or no reason to go deep |
| High share of neutral answers | Avoidance behavior, common on sensitive topics |
| Positive eNPS, negative town hall | The format is filtering critical voices out |
Concrete ways out
- One scale item per topic followed by a narrative question. That opens narrative space without bloating the questionnaire.
- AI voice interviews as a second format. Anonymous, no hierarchical interface, with follow-up questions.
- Establish credible anonymity. Never store audio, communicate minimum reporting thresholds clearly.
- Guarantee visible follow-up. Which three themes, which actions, which date.
- Add reverse-coded items to detect and correct acquiescence bias.
Why voice brings depth back
Speaking is less effortful than writing and at the same time richer in information. An AI that listens and asks targeted follow-up questions surfaces reasoning that no free-text field would capture. At the same time, the hierarchical interface disappears, because nobody on the other side is judging.
Frequently asked questions
Are high agreement scores automatically suspicious?
Not automatically, but when all items score equally high and free texts are missing, acquiescence bias is a plausible explanation.
How many employees does an anonymous voice interview need?
Common minimum thresholds are 5 to 10 answers per reporting filter. Smaller units aggregate into higher-level clusters.
Do we still need scales at all?
Yes, for comparability over time and across units. But scales only deliver the what, not the why. The combination is what creates a robust basis.
