ACM MUM 2024, first author

Voice to Value

You talk about your lunch. The app turns what you said into a review you would actually post, and keeps rewriting it until you say it is right.

Published
ACM MUM 2024, Stockholm. Read it (opens in a new tab)
Role
First author. Built the app, ran the study
Context
Master’s thesis, University of Oulu

01 The problem

Almost everyone has an opinion about lunch. Hardly anyone types it into a review form on a phone, standing up, after the moment has passed.

  • Typing takes effort
  • Speech comes out messy
  • AI-written reviews get trusted less

So the brief was narrow on purpose. Let people ramble, clean up only what speech gets wrong, and leave every change after that to them. The model edits. It does not write the review for you.

Comic strip in five panels. Voice-only: someone says “The uhh.. soup was cool” into a phone and the review reads exactly that. Assisted: someone says “Um… the cheese was smelly!”, asks the app to make it less rude, and gets “The cheese I ordered had a very strong odour. I recommend choosing carefully.”
The idea as drawn for the paper: one spoken review, without and with the assistant.

02 Cleaning up speech

Voice review

Uh yeah so I went to the restaurant and I had the um I don’t remember the name a hamburger with the wedges it was very good but maybe a bit too salty in the chips but yeah it was good.

Improved review

I visited the restaurant recently and tried their hamburger with wedges. The meal was quite enjoyable, although the chips were a tad too salty for my taste. Overall, it was a good dining experience.

Same meal, same verdict, same complaint about the salt. The cleanup runs at temperature 0.2, the least inventive setting in the app, and is told to keep the content and the tone.

03 One review, start to finish

“The soup was nice, but the bread was dry and the music was a bit too loud.”

  1. 01

    Both versions

    Speak

    Pick the restaurant and talk. A waveform moves while you do. The countdown only shows after 30 seconds, so nobody feels rushed, and there is room for five minutes.

  2. 02

    Both versions

    Transcribe

    OpenAI Whisper turns the recording into text straight away. In the voice-only version this is where it ends: fix it by hand, or submit.

  3. 03

    Assisted version

    Clean up

    GPT-4 removes the ums, fixes the grammar and drops anything off topic, keeping the content and the tone. The original transcript stays on screen next to it.

  4. 04

    Assisted version

    Ask for changes

    Type what should change and the agent rewrites the review. Repeat until it is right, or start again from the recording.

    Example

    “can you make it sound more positive and constructive”

  5. 05

    Assisted version

    Get ideas

    One tap for advice built on research into what makes a review useful: specific, detailed and easy to read.

    Example

    “the soup was very tasty”

    which soup, and what made it good?

  6. 06

    Both versions

    Rate and submit

    Quick sliders before sending: how willing you would be to share it, how happy you are with it, and whether the agent helped.

04 The app

Record, refine, rate

05 The study

Two versions, fourteen people, their own lunches

participants
14
reviews
157
reworked with the agent
42
requests for ideas
20
  1. Group A 8 people

    1. Voice only 5 reviews
    2. Assisted 5 reviews
  2. Group B 6 people

    1. Assisted 5 reviews
    2. Voice only 5 reviews

Real meals at campus restaurants, at least five reviews with each version, a short questionnaire after each and one at the end. The two groups started with different versions, so neither version got the advantage of going second.

06 What changed

Confidence nearly doubled

“How certain are you that you can leave a good review?” Out of 10.

  1. 4.86 Unaided
  2. 8.25 Voice only
  3. 9.00 Assisted
  4. 9.14 A tool like this
  • +31%

    Willingness to share

    4.53 → 5.93 out of 7, voice only against assisted. Significant, p < 0.05.

  • 6.15

    Satisfaction

    With the final review, out of 7, over 82 reviews.

  • 20/20

    Ideas marked helpful

    Every improvement tip that was asked for.

User experience scored higher for the assisted version on every dimension, 6.17 against 5.52, but none of those differences was statistically significant. And the unaided confidence score was asked at the end of the study, not the start, which the paper names as a limitation.

07 What people asked for

What people asked the agent to do

Content

  1. Add details 30
  2. Clarify ambiguity 10
  3. Omit information 8
  4. Correct information 5
  5. Adjust focus 2

Sentiment

  1. Less formal 6
  2. More excited 4
  3. More formal 3
  4. More positive 2
  5. Less excited 2
  6. Less positive 1

Style

  1. Easier to read 5
  2. Fit a platform 2
  3. Change the length 1

The most common request was for more, not less: people wanted their own review to say more of what they had noticed. The more reviews someone had written before, the longer their instructions were (Spearman 0.94).

08 In their words

“You can just explain the experience out loud as you would to a friend, and then you can just fix it up with the AI features. Genius.”
P17

Every participant preferred the assisted version.

What worked

  • “It is easier to leave a review as you can say the review in an unstructured way and the AI makes it structured.”
    P11
  • “When I use this app, it motivates me to do a review. I don’t need to think about grammar mistakes, and the AI-given review is very good.”
    P14
  • “people with broken English can create well written reviews”
    P6

What worried them

  • “if AI corrects a lot of the sentences, that will hide the people’s honest feelings”
    P12
  • “it suppresses our critical thinking ability.”
    P4
  • “people find it too easy to write restaurant reviews that they would say malicious things without thinking, making it easier to defame restaurants without a justified cause.”
    P9

09 What I built

The same project was my master’s thesis at the University of Oulu. The paper is the study written up for a wider audience.

What I built

Capture

A mobile web app: pick the restaurant, speak, watch the waveform move. Up to five minutes of recording, transcribed by Whisper.

A cleanup that keeps it yours

GPT-4 at a low temperature, told to keep the content and the tone and to remove only the fillers, the grammar slips and whatever went off topic.

The agent and the tips

Three separate model steps, each with its own prompt and temperature: the cleanup at 0.2, the agent at 0.8, and research-based tips at 1.

The field study

Fourteen people, two versions of the app, 157 reviews. Willingness to share the review went up 31% with the assistant, and confidence in writing a good one went up 88%. The write-up became my first-author paper at ACM MUM 2024. Read it (opens in a new tab)

  • HTML
  • Node.js
  • Express
  • OpenAI Whisper
  • GPT-4
  • Prompt design
  • Field study

Where it was cited

10 The paper

From Voice to Value: Leveraging AI to Enhance Spoken Online Reviews on the Go

Kavindu Perera, Dániel Szabó, Niels van Berkel, Aku Visuri, Chi-Lan Yang, Koji Yatani and Simo Hosio. University of Oulu, Aalborg University and the University of Tokyo.

ACM International Conference on Mobile and Ubiquitous Multimedia, Stockholm, 2024.

Read the paper (opens in a new tab)