Using AI at work

Map Free-Text Feedback into Themes with AI: Handle Multiple Labels and Unclassified Responses

This exercise classifies six short free-text comments into predefined themes and allows multiple labels when one response contains more than one topic. It also explains why the total number of labels can exceed the number of respondents and how to avoid unsupported conclusions about satisfaction or emotional intensity.

Show contents

Who this is forBeginners who want to organize short survey or review comments by theme but are unsure how much they can trust the counts and interpretation

What you need
  • You only need to be able to separate short comments into individual lines.
  • All F1–F6 responses in this exercise are fictional.
  • You do not need to enter personal data from real customers, employees, or participants.

01Prepare the six comments exactly as written with IDs

When classifying free-text feedback, first preserve each sentence as written instead of summarizing it immediately. A single short response can contain two different themes. In this example, F3 mentions both text size and insufficient location signage, so choosing only one would lose information. F6, which says no comment, should not be forced into a positive or negative category.

Sample input
F1 The information text is too small
F2 The location guidance is difficult to follow
F3 The information text is too small and there are not enough location signs
F4 The facilitator's explanation is friendly
F5 I would like to participate again
F6 No comment

The labels used in this exercise are visual guidance, location guidance, facilitator, repeat participation, and hold for classification. Emotional scores or satisfaction levels are not present in the source, so they should not be created.

Response IDKey expressionPossible number of labels
F1Information text is too small1
F2Location guidance is difficult1
F3Information text is too small and location signs are insufficient2
F4Facilitator's explanation is friendly1
F5Would like to participate again1
F6No comment1

02State the rules for multiple labels and hold for classification

When asking AI to classify themes, define not only the label names but also whether multiple labels are allowed. If one person's response can cover several themes, the total number of labels does not have to equal the number of respondents. In this example, the target counts are visual guidance 2, location guidance 2, facilitator 1, repeat participation 1, and hold for classification 1, for a total of 7 labels.

Prompt
Classify comments F1–F6 using the 5 labels below.

Labels:
- Visual guidance
- Location guidance
- Facilitator
- Repeat participation
- Hold for classification

Rules:
- Allow multiple labels when one comment clearly contains different themes.
- If two themes are explicitly present, as in F3, assign both labels.
- If a response has no evaluative content, such as no comment, assign hold for classification.
- Do not infer positive or negative emotional intensity, satisfaction scores, or whether all participants were satisfied.
- Do not add reasons that are not in the original comment.
- Preserve each response ID and original text.
- At the end, count the number of responses assigned to each label.

If the instruction allowing multiple labels is omitted, F3 may be forced into only one category. Allowing a hold category in advance also reduces the temptation to force an empty response into an unrelated theme.

03Check individual labels and totals together

The result below is not a measurement of the classification accuracy of any particular AI system. It is an editorial example created and reviewed with AI assistance to show how the input connects to label counts. A person should read F1–F6 one by one and verify the assigned labels.

Example result
F1 → Visual guidance
F2 → Location guidance
F3 → Visual guidance, Location guidance
F4 → Facilitator
F5 → Repeat participation
F6 → Hold for classification

Label counts:
Visual guidance 2
Location guidance 2
Facilitator 1
Repeat participation 1
Hold for classification 1

Number of respondents: 6
Total labels: 7
Because F3 has multiple labels, a total of 7 labels being greater than 6 respondents is not an error.
IDOriginal commentLabelContribution to count
F1The information text is too smallVisual guidanceVisual guidance +1
F2The location guidance is difficult to followLocation guidanceLocation guidance +1
F3The information text is too small and there are not enough location signsVisual guidance, Location guidanceVisual guidance +1, Location guidance +1
F4The facilitator's explanation is friendlyFacilitatorFacilitator +1
F5I would like to participate againRepeat participationRepeat participation +1
F6No commentHold for classificationHold for classification +1

04Do not expand theme classification into unsupported satisfaction analysis

A theme map organizes what topics were mentioned. It does not automatically determine participants' emotions or overall satisfaction. F4 says the facilitator was friendly and F5 says the respondent would like to participate again, but that does not support a claim that all 6 respondents were satisfied. The complaints in F1–F3 also do not provide a basis for assigning an intensity score.

Flawed result
Overall participant satisfaction was very high. All 6 participants were positive about the event, and F1–F3 only expressed mild dissatisfaction. F6 had no particular complaints, so that respondent can also be considered satisfied.
Unsupported inferenceWhy it is a problemWhat to record instead
All 6 were positiveF1–F3 and F6 do not provide evidence for overall satisfactionRecord only the theme labels for each response
Mild dissatisfactionEmotional intensity is not stated in the sourceRecord only the content about text size or location guidance
F6 was satisfiedNo comment does not mean satisfiedHold for classification

To reduce this type of error, treat themes, emotions, satisfaction, and behavioral intention as different analysis dimensions. This exercise classifies only themes. F5's statement about participating again is recorded only as repeat participation.

05Correct missing labels or unsupported interpretation with a targeted re-request

If F3 is classified only as location guidance or F6 is classified as positive, you can specify those IDs and ask for a correction. Providing expected totals as a check value also makes it easier for a person to verify the revised result.

Prompt
Review the classification result again.

- F3 contains both information text is too small and location signs are insufficient, so assign both Visual guidance and Location guidance.
- F6 says no comment and contains no evaluative content, so assign Hold for classification.
- Do not infer emotional intensity or satisfaction.

After correction, check whether the counts match the following:
Visual guidance 2
Location guidance 2
Facilitator 1
Repeat participation 1
Hold for classification 1

There are 6 respondents, but because multiple labels are allowed, the total number of labels is 7. Do not treat this difference as an error.

Providing expected totals does not mean classifications should be forced to match the numbers. The totals are check values that should be recalculated from the source. The final decision should come from verifying whether each response matches the label definition.

06Final check: keep respondent count and label count separate

With multiple labels, it is normal for the number of people and the number of labels to differ. When reporting results, do not treat 6 respondents and 7 labels as the same type of count. Use the checklist below to verify both the individual assignments and the totals.

Checklist
[ ] Is F1 labeled Visual guidance?
[ ] Is F2 labeled Location guidance?
[ ] Does F3 have both Visual guidance and Location guidance?
[ ] Is F4 labeled Facilitator?
[ ] Is F5 labeled Repeat participation?
[ ] Is F6 labeled Hold for classification?
[ ] Are the counts Visual guidance 2, Location guidance 2, Facilitator 1, Repeat participation 1, Hold for classification 1?
[ ] Are there 6 respondents and 7 total labels?
[ ] Is the total of 7 labels correctly treated as not an error?
[ ] Were unsupported interpretations such as everyone was satisfied or emotional intensity avoided?

With real free-text data, it is useful to test label definitions on a small set of examples before fixing the classification rules. This tutorial is an editorial exercise using 6 fictional comments and does not evaluate the representativeness or statistical significance of survey results.

What to check yourself

Editorial example in which a human recalculated the multiple-label assignments and label totals for the fictional F1–F6 free-text comments

  • Check that F3 has both Visual guidance and Location guidance
  • Check that F6 is Hold for classification
  • Check that the label counts are 2, 2, 1, 1, 1
  • Check that 6 respondents and 7 total labels are treated as different counts
  • Check that emotional intensity or universal satisfaction was not inferred
Verification limits

The result is an educational editorial example created and reviewed with AI assistance. It is not the result of analyzing real survey data, measuring classification accuracy, or verifying the automatic classification performance of any specific AI product.