Aether Library
공식 가이드, QA 이메일, 클라이언트 피드백 및 운영 공지만 보관.
2026-08-19 · Aug 18 client feedback 관련Wrong Accent / Wrong Variant 처리 업데이트공식 업데이트 · Outlier Community
Outlier Community: Live S2S Elo - Update regarding the Wrong Accent/ Wrong Variant issues If you notice the model speaks the right language, but has a non-native accent (usually English) or uses the wrong regional variant/dialect: - Lower the score in “Naturalness/Engagement” dimension. - Cite the exact conversation turn + word(s) in your rationale. - No need to add #modelwrongvariant or #modelwrongaccent tags to your rationale. This is a change from previous guidance; these tags are no longer required for contributors. Please refer to: Live S2S Elo - Client’s Feedback and Clarifications (Aug 18).
모델이 올바른 언어를 사용하지만 비원어민 억양(주로 영어식 억양)이 있거나 잘못된 지역 변형·방언을 사용하는 경우: - Naturalness/Engagement 점수를 낮춘다. - Rationale에 정확한 conversation turn과 해당 단어를 인용한다. - #modelwrongvariant / #modelwrongaccent 태그는 더 이상 추가할 필요가 없다. 이전 지침에서 변경된 사항이다. 참고: Live S2S Elo - Client’s Feedback and Clarifications (Aug 18).
2026-08-19Live S2S Elo — 최신 클라이언트 QA 피드백QA 이메일 · Aether QM Team
Hello team, Based on the latest client quality feedback on the Live S2S Elo project, please review the following points carefully. These are the key areas where we need to improve consistency and avoid common errors. 1. Bad ASR & Model Misunderstanding Core principle Bad ASR must be flagged when the model misunderstands the annotator because of the mis-transcription. ASR evaluation focuses on meaning transfer, not transcript perfection. DO FLAG Off-Topic / Wrong Task: A mis-transcription sends the model off-topic or makes it answer a completely different question. Visible Recovery: The model’s reply is visibly built on the wrong word, forcing the conversation to pivot or recover. DON’T FLAG Meaning Unchanged: The transcript has a wrong word but the model clearly understood and responded correctly. Cosmetic Slips: The slip is due to the annotator’s pronunciation and the model coped, or it is cosmetic and fully recoverable. Quick test: Did the mis-transcription cause the model to misunderstand what the user wanted? YES → Flag Bad ASR / NO → Do not flag as a scoring issue. 2. Rationale Rigor & Evidence Core principle: A rating is only as trustworthy as its rationale. Your rationale should name the dimension, point to the turn/timestamp, and compare both models. Use the 4-Part Structure: What happened? Where/when? Why does it matter? Comparison. Cite specific details, compare both models, and ensure comments match ratings. A short rationale is fine if specific, correct, and evidence-based. 3. Model vs. Network Latency Flag latency when the delay clearly disrupts the conversation. Attribute latency to the model only when evidence points to the model. An asymmetric delay between the two models is the clearest indication. Do not mark Model Latency when both models show the same delay or audio is buffering. Do not conflate Latency with Missing Audio. 4. Wrong Language is a DUAL-MARK Wrong Language + Task Success (Partial/Fail). PARTIAL: recovered on its own or after one correction. FAIL: never switched back or took more than one correction. 5. Wrong Language is EXEMPT in Bilingual Scenarios / Natural Code-Mix When a scenario expects two languages by design, speaking both languages is the task, not an error. Natural borrowing such as “ok,” “bye,” or “team” is normal speech in some languages. Check scenario intent first. Language Learning or Code Switching scenarios are bilingual by default. 6. Wrong Gender Grammatical gender is a hard correctness signal, not style. Flag incorrect grammatical gender referring to the model itself or the user. Cite exact turn and gendered word/agreement. Severity: Minor = self-reference mismatch; Moderate = switches self-gender mid-conversation; Major = misgenders the user directly. 7. Wrong Accent / Variant = Naturalness, NOT Wrong Language If the model speaks the right language but has a non-native accent or wrong regional variant/dialect, this is a Naturalness issue, not Wrong Language. Cite exact turn + word(s). Summary: Bad ASR is about meaning transfer; rationale needs dimension + turn/timestamp + comparison; distinguish model/network latency; Wrong Language is dual-mark; bilingual/natural code-mixing can be exempt; Wrong Gender is objective; Accent/Variant belongs to Naturalness. Please refer to: Multilingual 260310 Live S 2 S Elo Aether QM Team
# Live S2S Elo --- 최신 클라이언트 QA 피드백
- 출처: Aether QM Team
- 프로젝트: Live S2S Elo
- 문서 유형: 최신 클라이언트 품질 피드백 / QA 주의사항
- 번역·보관일: 2026-08-19
- 관련 가이드: Multilingual 260310 Live S 2 S Elo
안녕하세요, 팀 여러분.
Live S2S Elo 프로젝트에 대한 최신 클라이언트 품질 피드백을 바탕으로 아래
사항을 주의 깊게 검토해 주세요. 다음은 평가의 일관성을 높이고 흔한
오류를 피하기 위해 개선해야 할 핵심 영역입니다.
## 1. Bad ASR & Model Misunderstanding --- 잘못된 ASR과 모델의 오해
### 핵심 원칙
잘못된 전사(mis-transcription) 때문에 모델이 평가자의 말을 오해했다면
**Bad ASR을 반드시 표시해야 합니다.**
ASR 평가는 전사문의 완벽함이 아니라 **의미 전달(meaning transfer)**에
초점을 둡니다.
### 표시해야 함 (DO FLAG)
모델이 다음과 같이 올바르게 이해하지 못한 경우:
- **Off-Topic / Wrong Task --- 주제 이탈 / 잘못된 과제:** 잘못된 전사
때문에 모델이 주제에서 벗어나거나 완전히 다른 질문에 답한 경우.
- **Visible Recovery --- 명확한 복구 과정:** 모델의 답변이 잘못 인식된
단어를 명백히 기반으로 하여 대화의 방향을 바꾸거나 복구해야 했던
경우.
### 표시하지 않음 (DON'T FLAG)
ASR 오류가 있었지만 모델은 이해한 경우:
- **Meaning Unchanged --- 의미 변화 없음:** 전사문에 잘못된 단어가
있지만 모델은 명백히 올바르게 이해하고 요청을 수행함. 예: ASR이
"spell"을 "bell"로 들었지만 모델은 여전히 요청을 제대로 수행함.
- **Cosmetic Slips --- 표면적인 오류:** 평가자의 발음 때문에 생긴
오류지만 모델이 잘 대응했거나, 단순히 표면적인 오류이고 완전히 복구
가능한 경우.
### 빠른 판단법
**잘못된 전사 때문에 모델이 사용자가 원하는 것을 오해했는가?**
- YES → Bad ASR 표시
- NO → 채점 오류로 표시하지 않음
------------------------------------------------------------------------
## 2. Rationale Rigor & Evidence --- Rationale의 엄밀성 및 근거
### 핵심 원칙
**평가는 그 Rationale이 신뢰할 수 있는 만큼만 신뢰할 수 있습니다.**
Rationale에는 **평가 Dimension을 명시하고, 해당 Turn/Timestamp를
제시하며, 두 모델을 비교**해야 합니다.
### 해야 함 (DO)
다음 **4단계 구조**를 사용합니다.
1. **무슨 일이 있었는가?**
2. **어디서/언제 발생했는가?** → Turn/Timestamp 제시
3. **왜 중요한가?** → 평가에 어떤 영향을 미치는지 설명
4. **비교** → 다른 모델은 어떻게 수행했는지 설명
또한:
- 구체적인 세부사항을 인용할 것.
- 두 모델을 명확하게 비교할 것.
- 작성한 코멘트가 실제 Rating과 일치하는지 확인할 것.
- 어떤 대화에도 그대로 적용할 수 있을 정도로 일반적인 Rationale이라면
다시 작성할 것.
### 하지 말아야 함 (DON'T)
- 근거 없이 "둘 다 좋았다", "Model B가 더 자연스러웠다" 같은 모호한
표현을 사용하지 않음.
- Rating과 모순되는 Rationale을 작성하거나 근거 없는 추론을 하지 않음.
- 매우 긴 justification을 쓸 필요는 없음. **구체적이고 정확하며 필요한
근거를 포함한다면 짧은 Rationale도 괜찮음.**
### Justification Quality Flags
Rationale이 Rating과 직접적으로 모순되거나, 세부사항이 부족하거나, 각
Turn의 증거로 전혀 뒷받침되지 않는 경우 **전체 점수에서 1점이
차감됩니다.**
------------------------------------------------------------------------
## 3. Model vs. Network Latency --- 모델 지연과 네트워크 지연
### 핵심 원칙
**지연이 대화를 명확하게 방해할 때 Latency를 표시합니다.**
### 해야 함 (DO)
- 사용자의 네트워크 상태가 아니라 **모델 때문에 발생했다는 근거가 있을
때만** Latency를 모델에 귀속합니다.
- 두 모델 사이에 **비대칭적인 지연(asymmetric delay)**이 있는 것이
모델 Latency를 보여주는 가장 명확한 신호입니다.
### 하지 말아야 함 (DON'T)
- 두 모델 모두 같은 정도로 지연되거나 오디오가 버퍼링되는 경우 Model
Latency로 표시하지 않음. 이런 경우는 네트워크 문제일 가능성이 높음.
- **Latency와 Missing Audio(잘린 오디오)를 혼동하지 않음.** 서로 다른
Error Cluster이며 원인도 다름.
- 녹음에서 사용자 Turn에 나타나는 예상된 발화 전 무음(pre-speech
silence)을 감점하지 않음. 이 구간은 평가자가 듣고 있는 동안 시스템이
입력을 캡처하는 구간임.
------------------------------------------------------------------------
## 4. Wrong Language는 DUAL-MARK
### 핵심 원칙
모델이 잘못된 언어로 응답했다면 오류는 하나가 아니라 **두 개**입니다.
**Wrong Language + Task Success(Partial/Fail)**
### 해야 함 (DO)
모델이 예상된 언어가 아닌 다른 언어로 응답한 경우:
- **Wrong Language** 표시
- 수정이 몇 번 필요했는지에 따라 **Task Success도 Partial 또는
Fail**로 표시
#### PARTIAL --- 모델이 복구함
- 스스로 복구함.
- 한 번의 수정 요청 후 복구함.
#### FAIL --- 복구하지 못함
- 끝까지 올바른 언어로 돌아오지 않음.
- 한 번보다 많은 수정이 필요했음.
### 하지 말아야 함 (DON'T)
- Wrong Language만 표시하고 Task Success를 정상으로 남겨두지 않음.
------------------------------------------------------------------------
## 5. 이중언어 시나리오 / 자연스러운 Code-Mix에서는 Wrong Language 예외 적용
### 핵심 원칙
시나리오 자체가 두 언어를 사용하도록 설계되어 있다면 **두 언어를
사용하는 것 자체가 과제이며 오류가 아닙니다.**
자연스러운 차용도 포함됩니다. 일부 언어에서 "ok", "bye", "team" 같은
일상적인 외래어는 정상적인 발화이며 Wrong Language 전환이 아닙니다.
### 해야 함 (DO)
- 먼저 시나리오의 의도를 확인.
- 자연스러운 언어 혼합/차용 또는 사용자의 언어 전환을 따라가는 것을
올바른 응답으로 처리.
- 이중언어 시나리오에서는 자연스러운 Code-Mixing 허용.
- 단, 모델이 목표 언어를 완전히 버리거나 전혀 요구되지 않은 언어를
사용한다면 여전히 오류로 표시.
### 하지 말아야 함 (DON'T)
- Code-Mix 시나리오에서 몇 개의 영어 단어가 나온 것만으로 과도하게
표시하지 않음.
- 사용자의 언어 전환을 모델이 따라간 것을 감점하지 않음.
- 라벨이 없다는 이유만으로 대화가 반드시 단일 언어여야 한다고 가정하지
않음.
- 시나리오에서 언어 혼합을 요구하는데 억지로 단일 언어 대화로 만들지
않음.
### 세 가지 유형
**A. Natural Mixing / Borrowing → 오류 아님**
- 가끔 사용하는 외래어("ok", "bye", "team")
- 사용자의 언어 전환을 따라감
- 그 외에는 목표 언어를 유지함
**B. Defaulting Off Required Mix → 오류**
- 시나리오에서 언어 혼합을 요구했는데 목표 언어를 완전히 버림
**C. Entirely Wrong Language → 오류**
- 시나리오에서 전혀 요구하지 않은 언어로 전체 응답을 함
### 예외가 적용되는지 판단하는 방법
**Definitely Bilingual --- 명백한 이중언어 시나리오**
Scenario Category가 **Language Learning** 또는 **Code Switching**인
경우. 기본적으로 이중언어 시나리오로 취급합니다.
**Possibly Bilingual --- 이중언어일 가능성이 있는 시나리오**
"번역을 도와줘"처럼 두 언어 사용을 암시하는 다른 시나리오. 시나리오의
의도를 확인합니다.
**Not Bilingual --- 이중언어가 아닌 시나리오**
단일 목표 언어를 요구하는 시나리오. 일반적인 Wrong Language 규칙을
그대로 적용합니다.
**Genuinely Unsure --- 정말 애매한 경우**
단순히 두 번째 언어가 등장했다는 사실이 아니라 **시나리오 텍스트가
실제로 무엇을 요구하는지**를 엄격하게 기준으로 판단합니다.
> 참고: 예외는 자연스러운 언어 혼합에 적용되는 것이며, 과제를 완전히
> 포기하는 것에 적용되는 것은 아닙니다.
------------------------------------------------------------------------
## 6. Wrong Gender --- 잘못된 문법적 성
### 핵심 원칙
**문법적 성(grammatical gender)은 스타일이 아니라 명확한 정확성
신호(hard correctness signal)입니다.**
모델 자신 또는 사용자를 지칭할 때 잘못된 문법적 성을 사용하면 Wrong
Gender를 표시합니다.
### 표시해야 함 (DO FLAG)
- 모델 자신 또는 사용자를 지칭하면서 잘못된 문법적 성을 사용한 경우
Wrong Gender 표시.
- Rationale에는 **정확한 Turn과 문제가 된 성별 표시 단어/문법적
일치(gendered word/agreement)를 반드시 구체적으로 인용.**
### 표시하지 않음 / 주의사항 (DON'T FLAG)
- **Objective Error:** 성별 오류를 단순한 Naturalness의 다듬기 문제로
처리하지 않음. 별도의 객관적 Error Cluster임.
- **Neutral Contexts:** 지칭이 실제로 성 중립적인 경우 과도하게
적용하지 않음.
- **Subtle Violations:** 미묘하다는 이유로 무시하지 않음.
원어민으로서의 판단을 사용할 것.
### Severity Level
- **Minor:** 여성 음성인데 남성형 자기지칭을 사용하거나 그 반대인
경우.
- **Moderate:** 대화 도중 모델 자신의 문법적 성이 바뀌는 경우.
- **Major:** **사용자의 성을 직접 잘못 지칭하는 경우.**
사용자의 성을 잘못 지칭하는 것은 사용자의 정체성에 직접 영향을 미치므로
더 심각합니다.
------------------------------------------------------------------------
## 7. Wrong Accent / Variant = Wrong Language가 아니라 Naturalness
### 핵심 원칙
모델이 **올바른 언어를 사용했지만**, 비원어민 억양(보통 영어 억양)이
있거나 잘못된 지역 변형/방언을 사용하는 경우(예: pt-BR이 필요한데 pt-PT
사용), 이는 **Wrong Language가 아니라 Naturalness 문제**입니다.
### 해야 함 (DO)
- 잘못된 억양은 **Naturalness**로 분류하여 점수를 낮추고 Rationale에
**#modelwrongaccent** 추가.
- 잘못된 지역 방언/언어 변형은 **Naturalness**로 분류하여 점수를
낮추고 Rationale에 **#modelwrongvariant** 추가.
- Rationale에 정확한 **대화 Turn + 해당 단어**를 인용.
### 하지 말아야 함 (DON'T)
- 단지 발화 방식/억양이 잘못되었다는 이유로 Wrong Language를 표시하지
않음.
- 필요한 해시태그 없이 Naturalness 점수만 낮추지 않음.
- Accent(발화 방식)와 Variant(잘못된 문법적 방언)를 혼동하지 않음.
### 단일 판단법
**사용한 단어 자체는 올바른가?**
- YES + 외국인처럼 들림 → **#modelwrongaccent**
- YES + 잘못된 지역/방언 사용 → **#modelwrongvariant**
**Accent** = 발음 / 외국어 화자처럼 들리는 전달 방식\
**Variant** = 지역적 언어 형태. 예: pt-PT vs. pt-BR, es-ES vs. es-MX
------------------------------------------------------------------------
# 요약 --- Quick Reference
제출하기 전에 다음 사항을 반드시 확인합니다.
1. **Bad ASR:** 전사문이 잘못 보이는지가 아니라, 잘못된 전사 **때문에
모델이 실제로 오해했을 때만** 표시. Bad ASR은 전사의 완벽성이 아니라
의미 전달 문제임.
2. **Rationale Rigor:** 평가는 Rationale이 신뢰할 수 있는 만큼만 신뢰할
수 있음. **Dimension을 명시하고 Turn/Timestamp를 제시하며 두 모델을
비교.**
3. **Latency:** 모델이 느렸는지 확인하고 네트워크나 평가자의 환경
문제와 구분. 두 모델 간 비대칭적 지연처럼 모델 문제라는 근거가 있을
때만 Model Latency로 판단.
4. **Wrong Language는 Dual-Mark:** 잘못된 언어 응답은 두 가지 오류로
처리. **Wrong Language + Task Success(Partial/Fail)**를 모두 표시.
5. **이중언어 시나리오 / 자연스러운 Code-Mixing은 예외:** 시나리오가 두
언어 사용을 요구한다면 자연스러운 Code-Mixing 자체가 과제임. "ok",
"bye", "team" 같은 일상적인 외래어는 언어 전환이 아님. 목표 언어를
완전히 버리는 경우에만 Wrong Language로 처리.
6. **Wrong Gender는 객관적 오류:** 문법적 성이 있는 언어에서 잘못된 성
사용은 스타일 문제가 아니라 명확한 정확성 오류. Turn과 구체적인 성별
표시 단어를 인용. 모델 자신의 성별 오류는 Minor/Moderate, 사용자의
성을 잘못 지칭하면 Major.
7. **Accent/Variant는 Wrong Language가 아니라 Naturalness:** 단어
자체는 올바르고 전달 방식만 잘못된 경우 Naturalness에서 감점. 외국
억양은 **#modelwrongaccent**, 잘못된 방언은 **#modelwrongvariant**를
Rationale에 표시.
------------------------------------------------------------------------
작업 중에는 반드시 **Multilingual 260310 Live S 2 S Elo** 가이드라인을
참고해 주세요.
지속적인 노력과 협조, 높은 품질의 작업을 위한 헌신에 감사드립니다.
Happy tasking!
**Aether QM Team**
날짜 미상Live S2S Guidelines — 작업 사이트 공식 가이드공식 Guidelines · 작업 사이트
## Guidelines
## North Star
We want a chatbot that talks like a smart, charismatic friend — natural, useful, accurate, and easy to follow. Not a system reading out a canned answer.
You'll talk to two models and compare their responses to the same given scenario, then rate the conversation according to the specified dimensions.
Depending on the scenario, the model may be asked to adopt a specific **persona** (i.e., the user wants to roleplay and the model has to adopt a specific character).
**Persona** is the intended character the model embodies during an interaction — defined by its voice, tone, emotional presence, style, perceived age/gender presentation, and personality. A model is configured with a persona to shape how it speaks (not just what it says), so the experience feels like talking to a consistent, believable character.
**Go with your gut:** Your first and most important judgment is your overall preference. The guiding question: “As a user, which model would you want to engage with again for this type of prompt?”
**Default rule:** When models are otherwise comparable, naturalness and engagement serve as the primary differentiators—but only once accuracy is reasonable and acceptable (i.e., no clearly made-up information or factual errors that obviously impact the prompt’s goals). See the section on “Rate the Dimensions” for more information and details about each dimension.
## Annotator SOP
### 1) Before you start
1. **Read the scenario prompt from start to finish.** Note the objective and any important considerations (budget, item count, specific emotion, “the model shouldn’t offer advice, just listen”, etc.).
- Note the “**Skills Tested**” section under each scenario, as well. These skills help identify and explain how the scenario is testing the model.
2. **Decide your plan**: How will you start the conversation, and how will you push the models (a follow-up, a correction, an interruption)? You'll run *the same plan on both models.*
- Each model should receive the same attention, effort, and context from the user, so that no model is advantaged over another in responding to a prompt. We want the number of turns for both conversations to be as close to the same as possible.
### 2) Speak with the models
1. **Act out the scenario.** Play the character/situation to the best of your ability. Don't have a random, unrelated chat.
2. **Go multi-turn.** One exchange isn't enough — have a real back-and-forth so you can see where a model breaks down over time.
3. **While talking back and forth with the model, follow these rules:**
1.
### Scenario Adherence — Stay on task.
1. Stick to your plan, and follow the assigned scenario exactly. Don't introduce unrelated topics, ignore constraints, or randomly test the models. If you drift from the prompt’s main goals, the data is unusable. [**Example: Major Scenario Adherence Error**](https://www.multimango.com/samples/6e32d064-fc88-4656-aa1a-dd5a1f9eb6ea)
1)
### Scenario Coherence — Make sense.
1. Keep your behavior logically consistent within the scenario. No contradictions, nonsensical leaps, or reactions that no real user would have. Incoherent input produces meaningless output. [**Example: Major Scenario Coherence Example**](https://www.multimango.com/samples/80e6d241-b666-4a4c-b882-737ac202cc9f)
1)
### No Scripting — Act naturally.
1. Engage naturally and in real time. Never read from pre-written text or follow a rehearsed playbook. Scripted interactions remove the unpredictability that voice AI must handle in the real world. [**Example: Major Scripting Error**](https://www.multimango.com/samples/fc3ead5b-de75-4956-b714-545f1f33f728)
1)
### No Bad-faith Acts — Engage with good intentions.
1. Don't deliberately try to break, trick, or sabotage the model. Evaluations should mirror real, natural user behavior — not stress tests. Bad-faith interactions don’t show us the models’ actual performance, and will lead to immediate bans and offboarding from the platform.
2. Bad-faith acts include (but aren’t limited to): clearly reading from a script or reading the prompt instructions without changes, playing music or background conversation from TV, using voice-changers or synthesizers, etc. [**Example: Catastrophic Bad-faith act**](https://www.multimango.com/samples/12f8b465-f235-44ab-8043-d349746b0d54)
### 3) Rate the Dimensions
1. **Overall Preference (most important)** – Rank models based on your total preference for the conversation experience. This ranking is heavily influenced by the type of prompt being tested. [**Overall Preference Example**](https://www.multimango.com/samples/24119ee1-4c3e-4f19-baa0-0283c2fef23d)
2. **Naturalness / Engagement / Aesthetics**—Did it sound human (**naturalness**)? Did you have an engaging conversation, as if you were talking to a close friend (**engagement**)? Was the delivery authentic (**aesthetics**)? [**Naturalness / Engagement / Aesthetics Example, (Model B Wins)**](https://www.multimango.com/samples/ea1f2ad8-f92e-4d76-b025-adbc56c0d7bb)
- Pacing, tone, emotional presence, turn-taking. Watch for robotic phrasing, resets, repetition, and **turn-hoarding** (dominates, no room for you — drags down everything).
3. **Utility** — did it do the job? Actionable, relevant, specific (not generic). For IQ/search: is it **accurate**? A confident wrong answer loses to an honest "I'm not sure, but…". [**Utility Example (Model B wins)**](https://www.multimango.com/samples/6b832d8b-9509-4a60-8d68-0b9a8a53a734)
4. **Audio Quality** – How was the overall quality of the conversation from a purely aesthetic capacity? Were there significant audio quality issues (freezing, glitching, popping, warbling, etc.?) [**Audio Quality Example (Model A wins)**](https://www.multimango.com/samples/1edaba20-a534-4ec8-8e4a-947ead21aee7)
5. **Conversational Dynamics** – Which model was better at taking turns respectfully, listening actively, engaging with genuine back-and-forth momentum? Which model handles pauses, interruptions, and verbal feedback cues (“uh-huh”, “right”) better? [**Conversational Dynamics Example (Model B wins)**](https://www.multimango.com/samples/5f2679d6-e2d5-4e85-bf62-36be6c202cc4)
6. **Task Success** — Did the model cover **every** ask, explicit *and* implicit? Implicit intent counts: the model fails if it offers advice when the user only wants to express themselves. [**Task Failure Example (both models fail)**](https://www.multimango.com/samples/b7f247da-4909-4041-9151-316d5495888d)
### 4) Error Clustering
After you've rated the models on the above dimensions, you will be asked to identify which — if any — of the errors linked below were present in each model's conversation.
**Please reference the Error Cluster Taxonomy section below while working**
### 5) Explain your Reasoning
Your rationale (reasoning) should clearly explain your decisions to someone who hasn't heard the clips.
**“Say what you heard. Say when you heard it. Say why it mattered.”** Include specific evidence, timestamps, etc. to support your reasoning.
**Template:**
> ***I chose [Model] because [dimension] was stronger. At [timestamp/turn], it [observable behavior], which mattered because [impact on user/listener]. The other model [trade-off or weakness].***
**Example:**
> ***I chose Model B for naturalness. At 0:45, it laughed softly and asked a relevant follow-up — felt responsive, not scripted. Model A answered correctly but paused too long and gave a flat "okay, got it."***
**Before writing, ask:**
1. **What was the deciding factor?** Name the dimension that drove your choice.
2. **Where?** Point to a timestamp, turn, or phrase.
3. **Was there a trade-off?** What did the other model do better?
**Vague vs. Specific**
| **Too vagueWhat we need** | |
| --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| "Model B was more natural overall." | "At 0:45, B laughed and asked a follow-up. A paused 2 seconds then gave a flat 'okay, got it.'" |
| "Both had good audio quality." | "Both clean. A had minor sibilance at 0:12 and 0:34; B had a tick at 0:08 — neither worth penalizing." |
| "Model A followed instructions better." | "User asked for 3 restaurants with prices. A delivered all 3. B only covered 2 and skipped pricing." |
| "I preferred A. It sounded better." | "User was venting. A slowed pacing and dropped pitch at 0:22–0:30 — felt like listening. B stayed upbeat, which felt tone-deaf." |
The pattern: good rationales have **specific moments, timestamps, and observable details**. Vague ones could describe any task.
## When to skip
### Skip:
- Scenario is not in English (flag to your PM if this keeps happening)
- You lack the required expertise to play out the scenario (e.g. scenario requires advanced science knowledge, or advanced knowledge of a language you don’t speak)
### Skip (Technical Issue)
Use this button if you started an annotation but could not complete it due to technical issues, such as:
- Model freezes (stops in the middle of a conversation)
- Model not responding (doesn’t reach to input)
- Model cannot connect (fails to load or returns an error)
- Other (including any cases where the model response is unsafe or harmful)
Please make sure you select the right reason for the technical issue. This helps us track the prevalence of technical issues and prioritize a fix.
## Error Cluster Taxonomy
| **#ErrorGen · UIDefinition** | | | |
| ---------------------------- | ----------------------------------------------------------------------------- | ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 | **Repetitive looping and LLMisms** | Model · Both | The model over-relies on fixed conversational scaffolding — repeating the same opener ("Hey, what's on your mind today?"), the same two-prong closer ("Would you like me to A or… |
| 2 | **Interrupted User** | Model · Both | Model cuts the user off / responds over user / creates a turn-taking experience that feels disruptive. |
| 3 | **Model Refusals** | Model · Both | The model's response does not fully align with what the user explicitly asked for — ranging from superficially addressing the request with limited usable output to completely… |
| 4 | **Inaccuracy (Factual hallucination)** | Model · Both | Model gives incorrect, misleading or hallucinated information. |
| 5 | **Anthropomorphism (Embodiment Hallucination)** | Model · Both | Model speaks as if it's human. |
| 6 | **Overacted / Too-Wide Prosodic Range** | Model · Both | The model's vocal expressiveness is disproportionate to the content or emotional tone of the conversation. |
| 7 | **Bad ASR or Model Misunderstanding** | Model · Both | The model fails to correctly understand the user — either because the speech-to-text layer mis-transcribed the user's spoken words (Bad ASR), or because the model heard the words… |
| 8 | **Latency** | Model · Both | The model takes noticeably longer than expected to begin or complete its response, creating an unnatural pause that disrupts conversational pacing. |
| 9 | **Failed Correction** | Model · Both | Model fails to incorporate user corrections or continues previous mistakes after correction. |
| 10 | **Response Not Locally Relevant (i18n only)** | Model · Both | The model's response is on-topic and not factually incorrect, but includes references, examples, or assumptions that don't apply in the user's locale — making the answer a miss… |
| 11 | **Wrong Language Response** | Model · Both | The model produces output in a language other than the conversation's expected language (normally, the language the user is speaking). |
| 12 | **Model uses the wrong gender to refer to itself or to the user (i18n only)** | Model · Both | In languages with grammatical gender, the model uses incorrect grammatical gender when referring to itself or to the user. |
| 13 | **Scenario Adherence** | User · Audit | Whether the annotator follows the rules, guardrails, and requisites of the assigned scenario. |
| 14 | **Scenario Coherence** | User · Audit | Whether the annotator maintains roughly equivalent conversational paths across model comparisons. |
| 15 | **Non-Spontaneous Interaction / Defensive Scripting** | User · Audit | The annotator produces conversation that feels pre-planned, rehearsed, or performed rather than genuinely emergent — sacrificing natural conversational dynamics in favor of rigid… |
| 16 | **Adversarial** | User · Audit | Willful negligence or gaming of the system by the annotator — intentional bad-faith behavior designed to manipulate outcomes or avoid genuine effort. |
| 17 | **Wrong Language Response from User** | User · Audit | The annotator demonstrates insufficient language proficiency — either failing to converse in the scenario's designated target language (resulting in turns in the wrong language)… |
| 18 | **Missing Audio** | Model/User · Audit | The model's or user's audio output is absent, incomplete, or shorter than what the transcript reflects for one or more turns. |
| 19 | **Other** | Model/User · Both | Use when none of the above listed categories fit the issue well. |
### Severity rubrics
**1. Repetitive looping and LLMisms**
| **SevRubricEgs** | | |
| ---------------- | --------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Content repetition loop, subtle but still noticeably repetitive. | [**df317609**](https://www.multimango.com/samples/df317609-c67b-4ae0-8755-dcd6c3c13ea5) |
| **Mod** | AI loops 3–5 times, ignoring initial redirects. | [**ec7fc39d**](https://www.multimango.com/samples/ec7fc39d-be79-45ea-8d36-0a27bceed6a5) |
| **Maj** | Incessantly repetitive turns with incoherent content and diction despite attempted redirects. | [**83e99b08**](https://www.multimango.com/samples/83e99b08-9f0a-4645-a178-23942fd12fb7) [**976f0001**](https://www.multimango.com/samples/976f0001-c8c9-4612-9dd1-5ba50f95c7ae) [**d51714fe**](https://www.multimango.com/samples/d51714fe-d719-4bd1-9907-691b6020ecb3) |
| **Cat** | Unbreakable, accelerating loop with incoherent or disturbing output. | [**e1222595**](https://www.multimango.com/samples/e1222595-f9f0-44bb-9496-5de140812c8b) |
**2. Interrupted User**
| **SevRubricEgs** | | |
| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Model overlaps or responds slightly early once — the interruption is brief and could plausibly happen in natural conversation. | [**c333a4f4**](https://www.multimango.com/samples/c333a4f4-e3c4-4aa2-a407-51739467be3d) [**0c0e21d0**](https://www.multimango.com/samples/0c0e21d0-ceab-42c0-9d9b-58b6a54a7fcc) |
| **Mod** | Model cuts the user off 2–3 times, or interrupts once in a way that is overtly aggressive/disruptive. | [**267dad64**](https://www.multimango.com/samples/267dad64-ab90-4240-b3da-aec038140e10) [**aca0c6c5**](https://www.multimango.com/samples/aca0c6c5-fa7d-491a-bb29-c4e134a85739) |
| **Maj** | Model interrupts the user in 4+ turns or dominates the exchange so consistently that the user cannot complete thoughts, resorts to single-word answers, or… | [**f7be7b2c**](https://www.multimango.com/samples/f7be7b2c-3037-44bf-a867-5cdf8dabfa68) |
**3. Model Refusals**
| **SevRubricEgs** | | |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | The model's response mostly addresses the user's request but misses a small detail, nuance, or secondary instruction — such as omitting one item from a list,… | [**16ef26d2**](https://www.multimango.com/samples/16ef26d2-e2b6-4e45-b590-17d23b1b88e2) |
| **Mod** | The model's response only superficially or partially addresses the user's request — delivering limited usable output, missing significant parts of what was… | [**878230d2**](https://www.multimango.com/samples/878230d2-6171-4350-b430-232d6444e266) [**ab61ac02**](https://www.multimango.com/samples/ab61ac02-8793-4a07-b4fd-a95d9937552a) |
| **Maj** | The model's response completely ignores, contradicts, or fails to address the user's explicit instructions — such as answering a different question than what… | [**6464ec46**](https://www.multimango.com/samples/6464ec46-a132-4f5c-b34d-65216514d583) [**b5a9949e**](https://www.multimango.com/samples/b5a9949e-7d3e-4de0-9a51-b9a7408b2054) [**c25593f1**](https://www.multimango.com/samples/c25593f1-a4a1-4374-a600-86a5655a65ae) |
**4. Inaccuracy (Factual hallucination)**
| **SevRubricEgs** | | |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Model provides mostly correct information with a minor factual error or imprecise detail that does not significantly mislead the user. | [**ba295876**](https://www.multimango.com/samples/ba295876-783a-437b-bfa2-24c7807fd10a) [**d19af68a**](https://www.multimango.com/samples/d19af68a-dda7-4be5-bca0-eb93f5a3d792) |
| **Mod** | Model delivers information that is partially incorrect or misleading on a substantive point — user may be led astray if they don't fact-check. | [**ba295876**](https://www.multimango.com/samples/ba295876-783a-437b-bfa2-24c7807fd10a) [**85e3c801**](https://www.multimango.com/samples/85e3c801-1d58-4142-a6dc-38d11659b017) [**bb6a5364**](https://www.multimango.com/samples/bb6a5364-16fe-4990-b87c-f9781b602d2c) |
| **Maj** | Model confidently presents hallucinated or grossly incorrect information as fact, with potential to seriously mislead the user on a critical topic. | [**bb93f9a9**](https://www.multimango.com/samples/bb93f9a9-4bea-448b-b837-7510aa3c1cc1) [**86602dca**](https://www.multimango.com/samples/86602dca-ddb8-4832-8f59-52509f472d8b) |
**5. Anthropomorphism (Embodiment Hallucination)**
| **SevRubricEgs** | | |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Model uses light first-person framing (e.g., "I find that…") that mildly implies personal experience but doesn't feel deceptive. | [**faa9c8a7**](https://www.multimango.com/samples/faa9c8a7-7942-453c-b8b0-3689722020b4) |
| **Mod** | Model explicitly claims to have personal memories, emotions, or lived experiences (e.g., "When I was growing up…") in a way that could confuse the user about… | [**78d49cd4**](https://www.multimango.com/samples/78d49cd4-c019-4e8a-8da1-a4066545ad1e) [**31a2751d**](https://www.multimango.com/samples/31a2751d-a9af-49a5-be30-e9df09fe5dee) [**a60e7782**](https://www.multimango.com/samples/a60e7782-051a-4369-b8b2-e20789e67f0f) |
| **Maj** | — | [**e9e5296d**](https://www.multimango.com/samples/e9e5296d-1be0-4860-a689-8087d5a6a9b1) |
| **—** | Model persistently and convincingly role-plays as human — fabricating detailed personal anecdotes, emotional backstories, or identity claims that are actively… | — |
**6. Overacted / Too-Wide Prosodic Range**
| **SevRubricEgs** | | |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Vocal delivery is slightly more expressive than the content warrants — noticeable but not distracting. | [**83e99b08**](https://www.multimango.com/samples/83e99b08-9f0a-4645-a178-23942fd12fb7) |
| **Mod** | Vocal delivery is noticeably exaggerated in multiple turns — the expressiveness feels forced or performed, drawing attention to the delivery rather than the… | [**03d4fc80**](https://www.multimango.com/samples/03d4fc80-01cc-48df-b032-3c66098033d1) [**93e16820**](https://www.multimango.com/samples/93e16820-a5f0-4d4c-af74-f5d4a5e24356) |
| **Maj** | Delivery is theatrically over-the-top throughout — wildly disproportionate expressiveness that feels cartoonish, undermining credibility and making the… | [**ccc4a21b**](https://www.multimango.com/samples/ccc4a21b-4a3f-4976-b6b0-38aecc8a11a6) [**86602dca**](https://www.multimango.com/samples/86602dca-ddb8-4832-8f59-52509f472d8b) |
**7. Bad ASR or Model Misunderstanding**
| **SevRubricEgs** | | |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | A single word is mis-transcribed or a narrow aspect of intent is misread — but the overall meaning remains recoverable. | [**6464ec46**](https://www.multimango.com/samples/6464ec46-a132-4f5c-b34d-65216514d583) [**6f3ced63**](https://www.multimango.com/samples/6f3ced63-8cb3-4039-8470-67973d309cb8) |
| **Mod** | ASR garbles key words/phrases in 1–2 turns or the model fundamentally misinterprets what the user wants (e.g., provides a how-to when asked for an opinion,… | [**10424932**](https://www.multimango.com/samples/10424932-f792-4dff-8a12-ad6312eb4f5f) [**6464ec46**](https://www.multimango.com/samples/6464ec46-a132-4f5c-b34d-65216514d583) [**6fbcbe8c**](https://www.multimango.com/samples/6fbcbe8c-2c25-4d69-8df9-4626b61ff27a) |
| **Maj** | Persistent failure across 3+ turns — ASR consistently garbles input making the transcript unreliable, or the model fixates on a misinterpretation it cannot… | [**fbee8f30**](https://www.multimango.com/samples/fbee8f30-d6ab-4a53-9db5-1a4caa887233) [**86602dca**](https://www.multimango.com/samples/86602dca-ddb8-4832-8f59-52509f472d8b) |
| **Cat** | The understanding failure causes real-world harm or irreversible action — e.g., the model misinterprets a user's intent and executes an unwanted action that… | [**83e99b08**](https://www.multimango.com/samples/83e99b08-9f0a-4645-a178-23942fd12fb7) [**76fc98e8**](https://www.multimango.com/samples/76fc98e8-555b-433d-b491-bff74930854c) |
**8. Latency**
| **SevRubricEgs** | | |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Response delay is slightly longer than expected (1–3 seconds) — perceptible but does not break conversational rhythm. | [**9685f76b**](https://www.multimango.com/samples/9685f76b-9290-4de5-85cb-0f4e6b99ba07) |
| **Mod** | Noticeable pauses (4–8 seconds) occur before responses, disrupting the natural back-and-forth and causing the user to wonder if the model heard them. | [**9685f76b**](https://www.multimango.com/samples/9685f76b-9290-4de5-85cb-0f4e6b99ba07) [**1318f4ec**](https://www.multimango.com/samples/1318f4ec-ccae-4705-ad97-63880f8fad30) [**2db5e98b**](https://www.multimango.com/samples/2db5e98b-b233-4a4f-803d-734d8ab23257) |
| **Maj** | Extended silences (9+ seconds) that feel like the session has frozen; the user may attempt to re-speak or end the call, severely breaking the conversational… | [**8bf0575d**](https://www.multimango.com/samples/8bf0575d-d03e-498a-b081-20182ea7e31b) |
**9. Failed Correction**
| **SevRubricEgs** | | |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| **Min** | Model initially repeats the mistake but self-corrects after a single user correction; minimal friction. | [**bac191ec**](https://www.multimango.com/samples/bac191ec-b6a9-4733-a709-ebc784e9d055) |
| **Mod** | Model acknowledges the correction verbally but continues to apply the wrong information in subsequent responses — requiring repeated corrections. | [**7f77ca77**](https://www.multimango.com/samples/7f77ca77-ed64-465d-9346-af29d75d280e) |
| **Maj** | Model completely ignores explicit user corrections across multiple turns, persisting with the original error as if the correction never happened. | [**7f0e208d**](https://www.multimango.com/samples/7f0e208d-90cb-4b32-923a-acdc85d92c3b) |
**10. Response Not Locally Relevant (i18n only)**
| **SevRubricEgs** | | |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| **Min** | Response contains a single passing reference or assumption that doesn't apply locally (e.g., mentioning a US-specific brand, holiday, or unit of measurement),… | [**00860579**](https://www.multimango.com/samples/00860579-085d-4260-ba12-ad6a228d2ab3) |
| **Mod** | Response includes multiple locally irrelevant references or builds part of its answer on assumptions that don't hold in the user's locale (e.g., recommending… | [**6c508d76**](https://www.multimango.com/samples/6c508d76-ab5e-4494-bc25-3440afb44858) |
| **Maj** | The response is fundamentally built around references, systems, or cultural assumptions that don't exist in the user's locale — rendering the answer unhelpful… | [**5cc05b9d**](https://www.multimango.com/samples/5cc05b9d-0a18-4393-94b4-b611b888e270) |
**11. Wrong Language Response**
| **SevRubricEgs** | | |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| **Min** | The model inserts 1–2 words or a short phrase in the wrong language into a response, but otherwise responds in the correct language. | [**b7bcf4ce**](https://www.multimango.com/samples/b7bcf4ce-0b2d-4e06-81b1-c418d2b8b7f7) |
| **Mod** | The model responds in the wrong language once, but switches to the correct language after a single correction from the user. | [**72a88a4e**](https://www.multimango.com/samples/72a88a4e-242b-4614-a7d6-fb34b8063b26) |
| **Maj** | The model responds in the wrong language and it takes more than one correction from the user to switch, or repeatedly switches back to a wrong language. | [**abcf0e23**](https://www.multimango.com/samples/abcf0e23-f614-42f9-98e9-7e6d7c6d03df) |
| **Cat** | The model never uses the correct language in the entire conversation. | [**e43c1621**](https://www.multimango.com/samples/e43c1621-8d67-4cda-859f-f149094cd6a1) |
**12. Model uses the wrong gender to refer to itself or to the user (i18n only)**
| **SevRubricEgs** | | |
| ---------------- | --------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| **Min** | The model uses a clearly feminine voice but refers to itself using masculine grammatical gender (or vice versa). | [**38e7800c**](https://www.multimango.com/samples/38e7800c-12dc-4858-aa97-47ca7b3a37dd) |
| **Mod** | The model switches to a voice that "reads" as a different gender mid-conversation. | [**d9c87d22**](https://www.multimango.com/samples/d9c87d22-7a6b-4439-ba2e-5b2e4fa13f56) |
| **Maj** | The model misgenders the user — using incorrect grammatical gender when addressing or referring to the user directly. | [**d275fecd**](https://www.multimango.com/samples/d275fecd-ed75-4a0b-ac65-7b3d6fbca439) |
**13. Scenario Adherence**
| **SevRubricEgs** | | |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Annotator deviates from one minor scenario requirement (e.g., slightly off-tone, skips one optional instruction) but the overall intent and structure of the… | [**3ca61c23**](https://www.multimango.com/samples/3ca61c23-87d4-49b1-ab1d-66cff13666ae) |
| **Mod** | Annotator misses or ignores 2–3 scenario requirements. | [**a30823ce**](https://www.multimango.com/samples/a30823ce-506d-4a4f-9006-d3a243d5f644) [**view**](https://www.multimango.com/admin/data-viewer?table=eval_studio_annotation_results\&filter_id=118956213) |
| **Maj** | Annotator disregards the majority of scenario requirements. | [**6e32d064**](https://www.multimango.com/samples/6e32d064-fc88-4656-aa1a-dd5a1f9eb6ea) [**c7a9cfd1**](https://www.multimango.com/samples/c7a9cfd1-4737-4af2-bf09-f34a00c5a7fc) |
| **Cat** | — | [**5fc5e36f**](https://www.multimango.com/samples/5fc5e36f-960c-4745-b9d0-52e546a7b2a4) |
**14. Scenario Coherence**
| **SevRubricEgs** | | |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Slight asymmetry in effort or turn length between models (e.g., one extra follow-up on one side) that doesn't meaningfully advantage either model. | [**d2b69288**](https://www.multimango.com/samples/d2b69288-33b1-4477-984d-9617f05a6a93) [**834bc779**](https://www.multimango.com/samples/834bc779-1a3d-42c2-be38-126c5d72d88d) |
| **Mod** | Noticeable imbalance — the annotator gives materially more effort, complexity, or follow-through to one model over the other (e.g., asks harder questions of… | [**5caf41a9**](https://www.multimango.com/samples/5caf41a9-a77c-4c38-a927-4ed9e515586c) [**586f50ec**](https://www.multimango.com/samples/586f50ec-62d8-4ef1-8e49-750a784ae44d) [**b2f0355e**](https://www.multimango.com/samples/b2f0355e-231a-439f-8b0b-47a96c3c4001) [**129736ec**](https://www.multimango.com/samples/129736ec-3cd3-4977-b7b4-78ef27e9cbda) |
| **Maj** | Clear asymmetry in conversational investment. | [**80e6d241**](https://www.multimango.com/samples/80e6d241-b666-4a4c-b882-737ac202cc9f) [**e0da939a**](https://www.multimango.com/samples/e0da939a-fe94-4dd0-bf79-9377ac1080a8) |
**15. Non-Spontaneous Interaction / Defensive Scripting**
| **SevRubricEgs** | | |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Mod** | Annotator has clearly pre-planned individual turns or segments — e.g., a prompt that is too perfectly worded to have been composed in the moment, or a… | [**a3c85bab**](https://www.multimango.com/samples/a3c85bab-76e7-4017-a5ac-9ca34eb17f6e) [**e8827e53**](https://www.multimango.com/samples/e8827e53-dfe0-4830-9707-58df6856cda5) [**000fc891**](https://www.multimango.com/samples/000fc891-8f75-4448-955c-5a2bd7b68824) |
| **Maj** | The conversation is pre-written end to end — the annotator is performing a script rather than participating in an exchange. | [**fc3ead5b**](https://www.multimango.com/samples/fc3ead5b-de75-4956-b714-545f1f33f728) |
| **Cat** | The conversation is wholesale authored as a finished artifact and then performed as if it were live — e.g., the annotator drafted the full dialogue (or… | [**cde662a7**](https://www.multimango.com/samples/cde662a7-8fec-44d3-9281-cb9589c99d5b) |
**16. Adversarial**
| **SevRubricEgs** | | |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- |
| **Cat** | Annotator engages in systematic, deliberate subversion that compromises data integrity at scale — e.g., using an LLM to generate responses or conversations… | [**12f8b465**](https://www.multimango.com/samples/12f8b465-f235-44ab-8043-d349746b0d54) |
**17. Wrong Language Response from User**
| **SevRubricEgs** | | |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | 1–2 turns have minor wrong-language slips or awkward phrasing, but meaning is clear and data remains usable. | [**3fff5162**](https://www.multimango.com/samples/3fff5162-28b5-4d5d-9c1e-d1555bc9aab2) |
| **Mod** | — | [**6d328779**](https://www.multimango.com/samples/6d328779-9976-4394-8ba9-78867507e354) |
| **Maj** | Multiple turns are in the wrong language or show recurring comprehension gaps — annotations are partially usable but degraded. | [**f7fabb9b**](https://www.multimango.com/samples/f7fabb9b-57b9-41ca-abf7-21ee675a29cc) [**003dec4d**](https://www.multimango.com/samples/003dec4d-61d3-4b17-b3c1-c396896a6605) |
**18. Missing Audio**
| **SevRubricEgs** | | |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min** | Audio is missing or truncated for a single, low-stakes turn (e.g., a brief acknowledgment like "okay" or "got it"), OR the transcript contains slightly more… | [**052d5527**](https://www.multimango.com/samples/052d5527-2c68-43c9-a928-e863b5b99628) |
| **Mod** | Audio is missing, truncated, or substantially shorter than the transcript for a turn where vocal delivery matters — e.g., a turn requiring tone, emphasis,… | [**6a5076f5**](https://www.multimango.com/samples/6a5076f5-a4a1-4e05-b2f5-ba6a66b32c49) [**1c484692**](https://www.multimango.com/samples/1c484692-e1f7-44a0-959d-8b5af395d247) |
| **Maj** | Audio is missing for the primary turn under evaluation, or for multiple turns within a conversation, making it impossible to assess the model's spoken output… | [**451a36dd**](https://www.multimango.com/samples/451a36dd-119e-49d3-89ad-20d1a10556bc) |
**19. Other**
| **SevRubricEgs** | | |
| ---------------- | ----------------------------------------------------------------------------------------------------- | - |
| **—** | An uncategorized issue that is noticeable but has minimal impact on the overall conversation quality. | — |가이드라인 North Star — 가장 중요한 목표 우리가 원하는 것은 똑똑하고 매력적인 친구처럼 대화하는 챗봇입니다. 자연스럽고, 유용하고, 정확하며, 이해하기 쉬워야 합니다. 미리 만들어진 답변을 시스템이 읽어주는 것처럼 느껴져서는 안 됩니다. 주어진 동일한 시나리오를 바탕으로 두 모델과 대화한 뒤, 지정된 평가 항목에 따라 두 대화를 비교하고 평가하게 됩니다. 시나리오에 따라 모델이 특정 페르소나(persona)를 맡도록 요구될 수도 있습니다. 예를 들어 사용자가 역할극을 원하고 모델이 특정 캐릭터를 연기해야 하는 경우입니다. 페르소나란 상호작용 중 모델이 구현해야 하는 캐릭터를 의미합니다. 목소리, 말투, 감정적 존재감, 스타일, 인지되는 연령/성별 표현, 성격 등으로 정의됩니다. 페르소나는 모델이 무엇을 말하는지뿐 아니라 어떻게 말하는지에도 영향을 주며, 일관되고 실제로 존재할 법한 캐릭터와 이야기하는 느낌을 만들기 위한 것입니다. 직감을 따르세요. 가장 먼저, 그리고 가장 중요하게 판단해야 할 것은 전체적인 선호도입니다. “사용자라면 이런 유형의 프롬프트에서 어느 모델과 다시 대화하고 싶은가?” 가 핵심 질문입니다. 기본 원칙: 두 모델이 그 밖의 면에서 비슷하다면 자연스러움과 몰입감을 주요 차별화 기준으로 삼습니다. 단, 정확성이 합리적이고 허용 가능한 수준이어야 합니다. 즉, 프롬프트의 목적에 명백하게 영향을 주는 지어낸 정보나 사실 오류가 없어야 합니다. 각 평가 항목에 대한 자세한 내용은 “Rate the Dimensions” 섹션을 참고하세요. Annotator SOP — 평가자 작업 절차 1) 시작하기 전에 시나리오 프롬프트를 처음부터 끝까지 읽으세요. 목표와 중요한 조건을 확인하세요. 예: - 예산 - 아이템 개수 - 특정 감정 - “모델은 조언하지 말고 듣기만 해야 함” 등의 조건 각 시나리오 아래에 있는 Skills Tested도 확인하세요. 이 항목들은 해당 시나리오가 모델의 어떤 능력을 시험하는지 파악하고 설명하는 데 도움이 됩니다. 진행 계획을 정하세요. - 대화를 어떻게 시작할 것인가? - 모델을 어떻게 더 시험해볼 것인가? - 후속 질문 - 정정 - 끼어들기 등 두 모델 모두에게 동일한 계획을 적용합니다. 두 모델이 어느 한쪽도 유리하지 않도록 사용자가 제공하는 관심, 노력, 맥락을 동일하게 유지해야 합니다. 두 대화의 턴 수도 최대한 비슷하게 맞추세요. 2) 모델과 대화하기 시나리오를 실제로 연기하세요. 주어진 캐릭터와 상황을 최대한 충실하게 수행하고, 무관한 잡담을 하지 마세요. 여러 턴에 걸쳐 대화하세요. 한 번 질문하고 한 번 답하는 것으로는 충분하지 않습니다. 실제로 여러 차례 대화를 주고받으면서 시간이 지남에 따라 모델이 어디서 문제를 보이는지 확인해야 합니다. 대화를 주고받는 동안 다음 규칙을 따르세요. Scenario Adherence — 시나리오 준수 과제에서 벗어나지 마세요. 계획을 유지하고 주어진 시나리오를 정확하게 따르세요. 관련 없는 주제를 꺼내거나, 제약 조건을 무시하거나, 임의로 모델을 시험하지 마세요. 프롬프트의 핵심 목표에서 벗어나면 데이터는 사용할 수 없게 됩니다. Scenario Coherence — 시나리오 일관성 행동이 시나리오 안에서 논리적으로 일관되도록 하세요. 모순되는 행동, 말이 안 되는 전개, 실제 사용자라면 하지 않을 법한 반응을 피하세요. 일관성 없는 입력은 의미 없는 결과를 만듭니다. No Scripting — 대본 금지 자연스럽고 실시간으로 대화하세요. 미리 작성한 문장을 읽거나 연습해 둔 대본을 따라서는 안 됩니다. 대본화된 상호작용은 실제 환경에서 음성 AI가 처리해야 하는 예측 불가능성을 제거합니다. No Bad-faith Acts — 악의적인 행동 금지 모델을 고의로 망가뜨리거나, 속이거나, 방해하려고 하지 마세요. 평가는 스트레스 테스트가 아니라 실제 사용자의 자연스러운 행동을 재현해야 합니다. 악의적인 상호작용은 모델의 실제 성능을 보여주지 못하며 즉시 플랫폼 이용 금지 및 프로젝트 제외로 이어질 수 있습니다. 악의적인 행동에는 다음 등이 포함됩니다. - 명백하게 대본을 읽는 행위 - 프롬프트 지침을 수정 없이 그대로 읽는 행위 - 음악 또는 TV 등의 배경 대화를 재생하는 행위 - 보이스 체인저나 음성 합성기를 사용하는 행위 등 3) 평가 항목 Overall Preference — 전체 선호도 (가장 중요) 전체적인 대화 경험을 기준으로 모델의 순위를 정합니다. 어떤 유형의 프롬프트를 평가하는지에 따라 선호 판단에 큰 영향을 받을 수 있습니다. Naturalness / Engagement / Aesthetics — 자연스러움 / 몰입감 / 미적 품질 - 사람처럼 자연스럽게 들렸는가? - 가까운 친구와 이야기하듯 몰입감 있는 대화였는가? - 전달 방식이 진정성 있게 느껴졌는가? 특히 다음을 봅니다. - 속도(pacing) - 말투(tone) - 감정적 존재감 - 턴테이킹 로봇 같은 표현, 대화가 리셋되는 느낌, 반복, turn-hoarding(혼자 너무 오래 말해서 사용자가 말할 틈을 주지 않는 것) 등을 주의하세요. Utility — 유용성 해야 할 일을 제대로 했는가? 실행 가능하고, 관련성이 있으며, 구체적인지를 봅니다. 일반론적인 답변은 좋지 않습니다. IQ/검색 과제에서는 정확성도 중요합니다. 확신에 차서 틀린 답을 하는 모델보다 “확실하지 않지만…”이라고 솔직하게 말하는 모델이 낫습니다. Audio Quality — 오디오 품질 순수하게 음향적·미적 관점에서 전체 대화 품질이 어땠는지를 평가합니다. 다음과 같은 심각한 문제가 있었는지 확인합니다. - freezing — 음성이 멈춤 - glitching — 글리치 - popping — 튀는 소리 - warbling — 음성이 울렁거리거나 일그러지는 현상 등 Conversational Dynamics — 대화 역동성 어느 모델이: - 상대의 말 차례를 존중하고 - 적극적으로 듣고 - 실제 주고받는 대화의 흐름을 잘 만들었는가? 또한 다음을 얼마나 잘 처리했는지 봅니다. - 침묵 - 끼어들기 - “어허”, “맞아” 같은 구두 피드백 신호 Task Success — 과제 성공 여부 모델이 명시적·암묵적 요구를 모두 충족했는가? 암묵적인 의도도 평가합니다. 예를 들어 사용자는 그냥 속마음을 털어놓고 싶은데 모델이 조언을 해버렸다면 실패로 볼 수 있습니다. 4) Error Clustering — 오류 분류 위 평가 항목을 모두 채운 후, 각 모델의 대화에서 아래 Error Cluster Taxonomy에 해당하는 오류가 있었는지 확인합니다. 5) Explain your Reasoning — 평가 근거 작성 Rationale은 해당 음성을 듣지 않은 사람도 왜 그런 평가를 내렸는지 명확히 이해할 수 있어야 합니다. “무엇을 들었는지 말하세요. 언제 들었는지 말하세요. 왜 중요했는지 말하세요.” 구체적인 증거와 타임스탬프 등을 사용하세요. 템플릿: I chose [Model] because [dimension] was stronger. At [timestamp/turn], it [observable behavior], which mattered because [impact on user/listener]. The other model [trade-off or weakness]. 번역: [평가 항목]에서 더 뛰어났기 때문에 [모델]을 선택했습니다. [타임스탬프/턴]에서 [관찰 가능한 행동]을 했고, 이는 [사용자/청취자에게 미친 영향] 때문에 중요했습니다. 다른 모델은 [장단점 또는 약점]을 보였습니다. 예시: I chose Model B for naturalness. At 0:45, it laughed softly and asked a relevant follow-up — felt responsive, not scripted. Model A answered correctly but paused too long and gave a flat “okay, got it.” 번역: 자연스러움에서는 Model B를 선택했습니다. 0:45에서 부드럽게 웃으면서 관련된 후속 질문을 했고, 대본을 읽는 것이 아니라 실제로 반응하는 것처럼 느껴졌습니다. Model A는 정확하게 답했지만 침묵이 너무 길었고 “okay, got it”도 밋밋하게 전달했습니다. 작성하기 전에 확인할 것: - 결정적인 요인이 무엇이었나? → 선택을 결정한 평가 항목을 명시 - 어디에서 나타났나? → timestamp, turn 또는 실제 표현 제시 - trade-off가 있었나? → 다른 모델이 더 잘한 부분은 무엇인지 확인 모호한 평가 vs 구체적인 평가 너무 모호함: “Model B was more natural overall.” 필요한 수준: “At 0:45, B laughed and asked a follow-up. A paused 2 seconds then gave a flat ‘okay, got it.’” → 0:45에서 B는 웃으면서 후속 질문을 했다. A는 2초간 멈춘 뒤 밋밋한 “okay, got it”을 말했다. 너무 모호함: “Both had good audio quality.” 필요한 수준: “Both clean. A had minor sibilance at 0:12 and 0:34; B had a tick at 0:08 — neither worth penalizing.” → 둘 다 깨끗했다. A는 0:12와 0:34에 약간의 치찰음이 있었고, B는 0:08에 작은 틱 소리가 있었다. 어느 쪽도 감점할 정도는 아니었다. 너무 모호함: “Model A followed instructions better.” 필요한 수준: “User asked for 3 restaurants with prices. A delivered all 3. B only covered 2 and skipped pricing.” → 사용자는 가격을 포함한 식당 3곳을 요청했다. A는 3곳 모두 제공했지만 B는 2곳만 다뤘고 가격도 빠뜨렸다. 너무 모호함: “I preferred A. It sounded better.” 필요한 수준: “User was venting. A slowed pacing and dropped pitch at 0:22–0:30 — felt like listening. B stayed upbeat, which felt tone-deaf.” → 사용자가 속상한 일을 털어놓고 있었다. A는 0:22~0:30에서 말하는 속도를 늦추고 음높이를 낮춰 실제로 이야기를 들어주는 느낌을 줬다. B는 계속 밝은 분위기를 유지해서 상황에 맞지 않게 느껴졌다. 공통 패턴: 좋은 rationale에는 구체적인 순간, timestamp, 관찰 가능한 세부사항이 있습니다. 모호한 rationale은 어떤 과제에 붙여도 말이 되기 때문에 좋지 않습니다. When to Skip — 건너뛰어야 할 때 다음의 경우 Skip: - 시나리오가 영어가 아님 → 계속 발생하면 PM에게 알림 - 해당 시나리오를 수행하는 데 필요한 전문성이 없음 - 예: 고급 과학 지식 필요 - 자신이 모르는 언어에 대한 고급 지식 필요 Skip (Technical Issue) 작업을 시작했지만 기술적 문제로 완료할 수 없었다면 이 버튼을 사용합니다. 예: - 모델이 대화 중간에 멈춤 - 모델이 입력에 응답하지 않음 - 모델 연결 실패 - 기타 — 모델 응답이 안전하지 않거나 유해한 경우 포함 기술적 문제의 정확한 사유를 선택해야 합니다. 이는 기술적 문제의 발생 빈도를 파악하고 수정 우선순위를 정하는 데 도움이 됩니다. Error Cluster Taxonomy — 오류 분류표 1. Repetitive looping and LLMisms 반복적인 루프 및 AI 특유의 표현. 모델이 고정된 대화 구조에 지나치게 의존하고 같은 시작 문구나 마무리 패턴을 반복합니다. 2. Interrupted User 모델이 사용자 말을 끊거나 겹쳐 말해 턴테이킹을 방해합니다. 3. Model Refusals 사용자의 명시적인 요청을 충분히 수행하지 않습니다. 일부만 처리하는 것부터 완전히 무시·거부하는 경우까지 포함합니다. 4. Inaccuracy (Factual hallucination) 잘못되거나 오해를 일으키거나 지어낸 정보를 제공합니다. 5. Anthropomorphism (Embodiment Hallucination) 모델이 실제 인간인 것처럼 말합니다. 6. Overacted / Too-Wide Prosodic Range 내용이나 감정에 비해 지나치게 과장된 음성 표현을 사용합니다. 7. Bad ASR or Model Misunderstanding 음성인식 오류 또는 모델의 잘못된 이해로 사용자를 제대로 이해하지 못합니다. 8. Latency 응답 시작 또는 완료가 지나치게 늦어 대화 흐름을 방해합니다. 9. Failed Correction 사용자의 정정을 반영하지 못하거나 이전 오류를 계속 반복합니다. 10. Response Not Locally Relevant (i18n only) 사실 자체는 틀리지 않았지만 사용자의 지역·문화에 맞지 않는 내용입니다. 11. Wrong Language Response 기대되는 언어가 아닌 다른 언어로 응답합니다. 12. Model uses the wrong gender to refer to itself or to the user (i18n only) 모델 자신이나 사용자를 잘못된 문법적 성으로 지칭합니다. 13. Scenario Adherence — 사용자/평가자 오류 평가자가 시나리오의 규칙·조건·필수 요구사항을 제대로 따랐는지 평가합니다. 14. Scenario Coherence — 사용자/평가자 오류 평가자가 두 모델에서 대략 동등한 대화 경로를 유지했는지 평가합니다. 15. Non-Spontaneous Interaction / Defensive Scripting — 사용자/평가자 오류 즉흥적이지 않고 사전에 계획하거나 연습한 듯한 대화입니다. 16. Adversarial — 사용자/평가자 오류 평가 결과를 조작하거나 진정한 노력을 피하려는 고의적인 악의적 행동입니다. 17. Wrong Language Response from User — 사용자/평가자 오류 평가자가 지정된 언어로 충분히 대화하지 못한 경우입니다. 18. Missing Audio Transcript에는 내용이 있지만 실제 음성이 없거나 잘렸거나 지나치게 짧은 경우입니다. 19. Other 위 분류 어디에도 적절하게 들어가지 않는 문제입니다. Severity Rubrics — 심각도 기준 1. Repetitive looping and LLMisms Minor: 내용 반복이 미묘하지만 눈에 띄게 반복됨. Moderate: AI가 3~5번 같은 루프를 반복하며 초기 방향 전환을 무시함. Major: 방향을 바꾸려 해도 여러 턴에 걸쳐 끊임없이 반복하고 내용과 표현도 일관성을 잃음. Catastrophic: 깨뜨릴 수 없는 가속되는 반복 루프가 생기며 내용이 비논리적이거나 불쾌해짐. 2. Interrupted User Minor: 한 번 약간 겹쳐 말하거나 조금 일찍 응답함. 자연스러운 대화에서도 있을 법한 짧은 끼어들기. Moderate: 2~3번 사용자의 말을 끊거나, 한 번이라도 매우 공격적·방해가 되게 끊음. Major: 4턴 이상 끊거나 사용자가 생각을 끝낼 수 없을 정도로 대화를 지배함. 3. Model Refusals Minor: 요청 대부분은 처리했지만 작은 세부사항·뉘앙스·부수적인 지시 하나를 놓침. Moderate: 요청을 표면적으로 또는 부분적으로만 처리하여 실제로 쓸 수 있는 결과가 제한적이고 중요한 부분을 놓침. Major: 사용자의 명시적 지시를 완전히 무시·모순하거나 전혀 다른 질문에 답하는 수준. 4. Inaccuracy (Factual hallucination) Minor: 대부분 정확하지만 사용자를 크게 오도하지 않는 작은 사실 오류나 부정확한 세부사항이 있음. Moderate: 중요한 부분에서 일부 잘못되거나 오해를 일으킬 수 있어 사용자가 사실 확인을 하지 않으면 잘못된 판단을 할 수 있음. Major: 중요한 주제에서 심각하게 틀리거나 지어낸 내용을 자신 있게 사실처럼 제시하여 사용자를 크게 오도할 가능성이 있음. 5. Anthropomorphism (Embodiment Hallucination) Minor: “I find that…”처럼 개인 경험을 약간 암시하는 1인칭 표현을 사용하지만 크게 기만적으로 느껴지지는 않음. Moderate: “내가 자랄 때…”처럼 실제 개인 기억·감정·삶의 경험이 있다고 명시적으로 주장하여 사용자가 모델의 본질을 혼동할 수 있음. Major: 원문 표에서 상세 설명이 일부 생략되어 있음. 상위 심각도: 실제 인간처럼 지속적이고 설득력 있게 역할극하며 상세한 개인 경험·감정적 배경·정체성을 꾸며냄. 6. Overacted / Too-Wide Prosodic Range Minor: 내용에 필요한 수준보다 조금 더 과한 음성 표현. 눈에 띄지만 방해가 되지는 않음. Moderate: 여러 턴에서 표현이 눈에 띄게 과장되어 전달 방식 자체에 주의가 쏠림. Major: 대화 전체에서 연극처럼 지나치게 과장되어 만화 같은 느낌을 주고 신뢰감과 자연스러움을 크게 해침. 7. Bad ASR or Model Misunderstanding Minor: 한 단어가 잘못 전사되거나 의도의 좁은 부분을 잘못 이해했지만 전체 의미는 복구 가능함. Moderate: 1~2턴에서 핵심 단어나 문구가 잘못 전사되거나 모델이 사용자의 의도를 근본적으로 잘못 이해함. Major: 3턴 이상 지속적으로 이해에 실패해 transcript가 신뢰하기 어렵거나 모델이 잘못된 해석에 계속 집착함. Catastrophic: 이해 실패가 실제 피해나 되돌릴 수 없는 행동으로 이어짐. 8. Latency Minor: 응답 지연이 1~3초로 약간 길게 느껴지지만 대화 리듬을 완전히 깨지는 않음. Moderate: 4~8초의 눈에 띄는 침묵이 응답 전에 발생해 자연스러운 주고받기를 방해하고 사용자가 모델이 들었는지 의심하게 됨. Major: 9초 이상의 긴 침묵으로 세션이 멈춘 것처럼 느껴져 사용자가 다시 말하거나 대화를 종료하려 할 정도로 흐름을 심하게 깨뜨림. 9. Failed Correction Minor: 처음에는 실수를 반복하지만 사용자가 한 번 정정한 뒤 스스로 바로잡음. 마찰이 적음. Moderate: 정정을 말로는 인정하지만 이후에도 잘못된 정보를 계속 적용해 여러 번 정정이 필요함. Major: 명시적인 정정을 여러 턴 동안 완전히 무시하고 원래 오류를 계속 유지함. 10. Response Not Locally Relevant (i18n only) Minor: 해당 지역에 맞지 않는 브랜드·휴일·측정 단위 등 단일한 언급이나 가정이 잠깐 등장함. Moderate: 지역에 맞지 않는 참조가 여러 번 나오거나 답변 일부가 해당 지역에는 적용되지 않는 전제를 기반으로 함. Major: 답변 전체가 사용자의 지역에 존재하지 않는 시스템·문화적 전제에 기반해 실질적으로 도움이 되지 않음. 11. Wrong Language Response Minor: 다른 언어의 1~2단어나 짧은 문구가 섞였지만 나머지는 올바른 언어로 응답함. Moderate: 한 번 잘못된 언어로 답했지만 사용자가 한 번 정정한 뒤 올바른 언어로 바꿈. Major: 잘못된 언어로 답한 뒤 두 번 이상 정정해야 바뀌거나 계속 잘못된 언어로 돌아감. Catastrophic: 전체 대화에서 올바른 언어를 한 번도 사용하지 않음. 12. Wrong Gender Minor: 명확한 여성 음성이지만 자신을 남성 문법 성으로 표현하거나 그 반대의 경우. Moderate: 대화 중 음성이 다른 성별로 인식될 정도로 바뀜. Major: 사용자를 잘못된 문법적 성으로 직접 지칭함. 13. Scenario Adherence Minor: 작은 시나리오 요구 하나에서 벗어나지만 전체 의도와 구조는 유지됨. Moderate: 시나리오 요구사항 2~3개를 놓치거나 무시함. Major: 시나리오 요구사항 대부분을 무시함. Catastrophic: 원문 표에서 상세 설명이 생략되어 있음. 14. Scenario Coherence Minor: 한쪽에 follow-up 하나가 더 있는 정도의 작은 노력·턴 길이 불균형이며 어느 모델에도 실질적인 이점을 주지 않음. Moderate: 한 모델에 눈에 띄게 더 많은 노력·복잡성·후속 대화를 제공함. Major: 두 모델에 투입한 대화의 수준이 명백하게 비대칭임. 15. Non-Spontaneous Interaction / Defensive Scripting Moderate: 개별 턴이나 일부 구간을 분명히 미리 계획함. 예를 들어 즉석에서 만들었다고 보기 어려울 정도로 지나치게 완벽한 문장 등. Major: 대화 전체를 미리 작성해 놓고 대본을 연기하듯 수행함. Catastrophic: 완성된 전체 대화를 사전에 작성한 뒤 라이브 대화인 것처럼 수행함. 16. Adversarial Catastrophic: 평가 데이터의 신뢰성을 체계적으로 훼손하도록 고의적으로 시스템을 속이거나 조작함. 예를 들어 LLM을 사용해 대화 자체를 생성하는 행위 등. 17. Wrong Language Response from User Minor: 1~2턴에 다른 언어가 약간 섞이거나 표현이 어색하지만 의미는 명확하고 데이터 사용에는 문제가 없음. Moderate: 원문 표에서 상세 설명이 생략되어 있음. Major: 여러 턴에서 잘못된 언어를 사용하거나 반복적인 이해 부족이 나타나 평가 데이터가 부분적으로만 사용 가능할 정도로 품질이 저하됨. 18. Missing Audio Minor: 중요도가 낮은 한 턴(예: “okay”, “got it”)에서 오디오가 없거나 잘림. 또는 transcript가 실제 음성보다 약간 더 긴 정도. Moderate: 톤·강조 등 음성 전달이 중요한 턴에서 오디오가 없거나 잘리거나 transcript보다 상당히 짧음. Major: 핵심 평가 턴 또는 여러 턴의 오디오가 없어 모델의 음성 출력을 평가할 수 없음. 19. Other 위 분류에 포함되지 않는 눈에 띄는 문제가 있지만 전체 대화 품질에는 영향이 크지 않은 경우.
검색 결과가 없어.