Aether Library

공식 가이드, QA 이메일, 클라이언트 피드백 및 운영 공지만 보관.

×
2026-08-19 · Aug 18 client feedback 관련
Wrong Accent / Wrong Variant 처리 업데이트
공식 업데이트 · Outlier Community
Outlier Community: Live S2S Elo - Update regarding the Wrong Accent/ Wrong Variant issues

If you notice the model speaks the right language, but has a non-native accent (usually English) or uses the wrong regional variant/dialect:
- Lower the score in “Naturalness/Engagement” dimension.
- Cite the exact conversation turn + word(s) in your rationale.
- No need to add #modelwrongvariant or #modelwrongaccent tags to your rationale. This is a change from previous guidance; these tags are no longer required for contributors.

Please refer to: Live S2S Elo - Client’s Feedback and Clarifications (Aug 18).
2026-08-19
Live S2S Elo — 최신 클라이언트 QA 피드백
QA 이메일 · Aether QM Team
Hello team,

Based on the latest client quality feedback on the Live S2S Elo project, please review the following points carefully. These are the key areas where we need to improve consistency and avoid common errors.

1. Bad ASR & Model Misunderstanding

Core principle
Bad ASR must be flagged when the model misunderstands the annotator because of the mis-transcription.
ASR evaluation focuses on meaning transfer, not transcript perfection.

DO FLAG
Off-Topic / Wrong Task: A mis-transcription sends the model off-topic or makes it answer a completely different question.
Visible Recovery: The model’s reply is visibly built on the wrong word, forcing the conversation to pivot or recover.

DON’T FLAG
Meaning Unchanged: The transcript has a wrong word but the model clearly understood and responded correctly.
Cosmetic Slips: The slip is due to the annotator’s pronunciation and the model coped, or it is cosmetic and fully recoverable.

Quick test: Did the mis-transcription cause the model to misunderstand what the user wanted? YES → Flag Bad ASR / NO → Do not flag as a scoring issue.

2. Rationale Rigor & Evidence
Core principle: A rating is only as trustworthy as its rationale. Your rationale should name the dimension, point to the turn/timestamp, and compare both models.
Use the 4-Part Structure: What happened? Where/when? Why does it matter? Comparison.
Cite specific details, compare both models, and ensure comments match ratings. A short rationale is fine if specific, correct, and evidence-based.

3. Model vs. Network Latency
Flag latency when the delay clearly disrupts the conversation. Attribute latency to the model only when evidence points to the model. An asymmetric delay between the two models is the clearest indication. Do not mark Model Latency when both models show the same delay or audio is buffering. Do not conflate Latency with Missing Audio.

4. Wrong Language is a DUAL-MARK
Wrong Language + Task Success (Partial/Fail). PARTIAL: recovered on its own or after one correction. FAIL: never switched back or took more than one correction.

5. Wrong Language is EXEMPT in Bilingual Scenarios / Natural Code-Mix
When a scenario expects two languages by design, speaking both languages is the task, not an error. Natural borrowing such as “ok,” “bye,” or “team” is normal speech in some languages. Check scenario intent first. Language Learning or Code Switching scenarios are bilingual by default.

6. Wrong Gender
Grammatical gender is a hard correctness signal, not style. Flag incorrect grammatical gender referring to the model itself or the user. Cite exact turn and gendered word/agreement. Severity: Minor = self-reference mismatch; Moderate = switches self-gender mid-conversation; Major = misgenders the user directly.

7. Wrong Accent / Variant = Naturalness, NOT Wrong Language
If the model speaks the right language but has a non-native accent or wrong regional variant/dialect, this is a Naturalness issue, not Wrong Language. Cite exact turn + word(s).

Summary: Bad ASR is about meaning transfer; rationale needs dimension + turn/timestamp + comparison; distinguish model/network latency; Wrong Language is dual-mark; bilingual/natural code-mixing can be exempt; Wrong Gender is objective; Accent/Variant belongs to Naturalness.

Please refer to: Multilingual 260310 Live S 2 S Elo

Aether QM Team
날짜 미상
Live S2S Guidelines — 작업 사이트 공식 가이드
공식 Guidelines · 작업 사이트
## Guidelines

## North Star

We want a chatbot that talks like a smart, charismatic friend — natural, useful, accurate, and easy to follow. Not a system reading out a canned answer.

You'll talk to two models and compare their responses to the same given scenario, then rate the conversation according to the specified dimensions.

Depending on the scenario, the model may be asked to adopt a specific **persona** (i.e., the user wants to roleplay and the model has to adopt a specific character).

**Persona** is the intended character the model embodies during an interaction — defined by its voice, tone, emotional presence, style, perceived age/gender presentation, and personality. A model is configured with a persona to shape how it speaks (not just what it says), so the experience feels like talking to a consistent, believable character.

**Go with your gut:** Your first and most important judgment is your overall preference. The guiding question: “As a user, which model would you want to engage with again for this type of prompt?”

**Default rule:** When models are otherwise comparable, naturalness and engagement serve as the primary differentiators—but only once accuracy is reasonable and acceptable (i.e., no clearly made-up information or factual errors that obviously impact the prompt’s goals). See the section on “Rate the Dimensions” for more information and details about each dimension.

## Annotator SOP

### 1) Before you start

1. **Read the scenario prompt from start to finish.** Note the objective and any important considerations (budget, item count, specific emotion, “the model shouldn’t offer advice, just listen”, etc.).
   - Note the “**Skills Tested**” section under each scenario, as well. These skills help identify and explain how the scenario is testing the model.
2. **Decide your plan**: How will you start the conversation, and how will you push the models (a follow-up, a correction, an interruption)? You'll run *the same plan on both models.*
   - Each model should receive the same attention, effort, and context from the user, so that no model is advantaged over another in responding to a prompt. We want the number of turns for both conversations to be as close to the same as possible.

### 2) Speak with the models

1. **Act out the scenario.** Play the character/situation to the best of your ability. Don't have a random, unrelated chat.
2. **Go multi-turn.** One exchange isn't enough — have a real back-and-forth so you can see where a model breaks down over time.
3. **While talking back and forth with the model, follow these rules:**
   1.
   ### Scenario Adherence — Stay on task.
   1. Stick to your plan, and follow the assigned scenario exactly. Don't introduce unrelated topics, ignore constraints, or randomly test the models. If you drift from the prompt’s main goals, the data is unusable. [**Example: Major Scenario Adherence Error**](https://www.multimango.com/samples/6e32d064-fc88-4656-aa1a-dd5a1f9eb6ea)
   1)
   ### Scenario Coherence — Make sense.
   1. Keep your behavior logically consistent within the scenario. No contradictions, nonsensical leaps, or reactions that no real user would have. Incoherent input produces meaningless output. [**Example: Major Scenario Coherence Example**](https://www.multimango.com/samples/80e6d241-b666-4a4c-b882-737ac202cc9f)
   1)
   ### No Scripting — Act naturally.
   1. Engage naturally and in real time. Never read from pre-written text or follow a rehearsed playbook. Scripted interactions remove the unpredictability that voice AI must handle in the real world. [**Example: Major Scripting Error**](https://www.multimango.com/samples/fc3ead5b-de75-4956-b714-545f1f33f728)
   1)
   ### No Bad-faith Acts — Engage with good intentions.
   1. Don't deliberately try to break, trick, or sabotage the model. Evaluations should mirror real, natural user behavior — not stress tests. Bad-faith interactions don’t show us the models’ actual performance, and will lead to immediate bans and offboarding from the platform.
   2. Bad-faith acts include (but aren’t limited to): clearly reading from a script or reading the prompt instructions without changes, playing music or background conversation from TV, using voice-changers or synthesizers, etc. [**Example: Catastrophic Bad-faith act**](https://www.multimango.com/samples/12f8b465-f235-44ab-8043-d349746b0d54)

### 3) Rate the Dimensions

1. **Overall Preference (most important)** – Rank models based on your total preference for the conversation experience. This ranking is heavily influenced by the type of prompt being tested. [**Overall Preference Example**](https://www.multimango.com/samples/24119ee1-4c3e-4f19-baa0-0283c2fef23d)
2. **Naturalness / Engagement / Aesthetics**—Did it sound human (**naturalness**)? Did you have an engaging conversation, as if you were talking to a close friend (**engagement**)? Was the delivery authentic (**aesthetics**)? [**Naturalness / Engagement / Aesthetics Example, (Model B Wins)**](https://www.multimango.com/samples/ea1f2ad8-f92e-4d76-b025-adbc56c0d7bb)

- Pacing, tone, emotional presence, turn-taking. Watch for robotic phrasing, resets, repetition, and **turn-hoarding** (dominates, no room for you — drags down everything).

3. **Utility** — did it do the job? Actionable, relevant, specific (not generic). For IQ/search: is it **accurate**? A confident wrong answer loses to an honest "I'm not sure, but…". [**Utility Example (Model B wins)**](https://www.multimango.com/samples/6b832d8b-9509-4a60-8d68-0b9a8a53a734)
4. **Audio Quality** – How was the overall quality of the conversation from a purely aesthetic capacity? Were there significant audio quality issues (freezing, glitching, popping, warbling, etc.?) [**Audio Quality Example (Model A wins)**](https://www.multimango.com/samples/1edaba20-a534-4ec8-8e4a-947ead21aee7)
5. **Conversational Dynamics** – Which model was better at taking turns respectfully, listening actively, engaging with genuine back-and-forth momentum? Which model handles pauses, interruptions, and verbal feedback cues (“uh-huh”, “right”) better? [**Conversational Dynamics Example (Model B wins)**](https://www.multimango.com/samples/5f2679d6-e2d5-4e85-bf62-36be6c202cc4)
6. **Task Success** — Did the model cover **every** ask, explicit *and* implicit? Implicit intent counts: the model fails if it offers advice when the user only wants to express themselves. [**Task Failure Example (both models fail)**](https://www.multimango.com/samples/b7f247da-4909-4041-9151-316d5495888d)

### 4) Error Clustering

After you've rated the models on the above dimensions, you will be asked to identify which — if any — of the errors linked below were present in each model's conversation.

**Please reference the Error Cluster Taxonomy section below while working**

### 5) Explain your Reasoning

Your rationale (reasoning) should clearly explain your decisions to someone who hasn't heard the clips.

**“Say what you heard. Say when you heard it. Say why it mattered.”** Include specific evidence, timestamps, etc. to support your reasoning.

**Template:**

> ***I chose [Model] because [dimension] was stronger. At [timestamp/turn], it [observable behavior], which mattered because [impact on user/listener]. The other model [trade-off or weakness].***

**Example:**

> ***I chose Model B for naturalness. At 0:45, it laughed softly and asked a relevant follow-up — felt responsive, not scripted. Model A answered correctly but paused too long and gave a flat "okay, got it."***

**Before writing, ask:**

1. **What was the deciding factor?** Name the dimension that drove your choice.
2. **Where?** Point to a timestamp, turn, or phrase.
3. **Was there a trade-off?** What did the other model do better?

**Vague vs. Specific**

| **Too vagueWhat we need**               |                                                                                                                                  |
| --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| "Model B was more natural overall."     | "At 0:45, B laughed and asked a follow-up. A paused 2 seconds then gave a flat 'okay, got it.'"                                  |
| "Both had good audio quality."          | "Both clean. A had minor sibilance at 0:12 and 0:34; B had a tick at 0:08 — neither worth penalizing."                           |
| "Model A followed instructions better." | "User asked for 3 restaurants with prices. A delivered all 3. B only covered 2 and skipped pricing."                             |
| "I preferred A. It sounded better."     | "User was venting. A slowed pacing and dropped pitch at 0:22–0:30 — felt like listening. B stayed upbeat, which felt tone-deaf." |

The pattern: good rationales have **specific moments, timestamps, and observable details**. Vague ones could describe any task.

## When to skip

### Skip:

- Scenario is not in English (flag to your PM if this keeps happening)
- You lack the required expertise to play out the scenario (e.g. scenario requires advanced science knowledge, or advanced knowledge of a language you don’t speak)

### Skip (Technical Issue)

Use this button if you started an annotation but could not complete it due to technical issues, such as:

- Model freezes (stops in the middle of a conversation)
- Model not responding (doesn’t reach to input)
- Model cannot connect (fails to load or returns an error)
- Other (including any cases where the model response is unsafe or harmful)

Please make sure you select the right reason for the technical issue. This helps us track the prevalence of technical issues and prioritize a fix.

## Error Cluster Taxonomy

| **#ErrorGen · UIDefinition** |                                                                               |                    |                                                                                                                                                                                     |
| ---------------------------- | ----------------------------------------------------------------------------- | ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1                            | **Repetitive looping and LLMisms**                                            | Model · Both       | The model over-relies on fixed conversational scaffolding — repeating the same opener ("Hey, what's on your mind today?"), the same two-prong closer ("Would you like me to A or…   |
| 2                            | **Interrupted User**                                                          | Model · Both       | Model cuts the user off / responds over user / creates a turn-taking experience that feels disruptive.                                                                              |
| 3                            | **Model Refusals**                                                            | Model · Both       | The model's response does not fully align with what the user explicitly asked for — ranging from superficially addressing the request with limited usable output to completely…     |
| 4                            | **Inaccuracy (Factual hallucination)**                                        | Model · Both       | Model gives incorrect, misleading or hallucinated information.                                                                                                                      |
| 5                            | **Anthropomorphism (Embodiment Hallucination)**                               | Model · Both       | Model speaks as if it's human.                                                                                                                                                      |
| 6                            | **Overacted / Too-Wide Prosodic Range**                                       | Model · Both       | The model's vocal expressiveness is disproportionate to the content or emotional tone of the conversation.                                                                          |
| 7                            | **Bad ASR or Model Misunderstanding**                                         | Model · Both       | The model fails to correctly understand the user — either because the speech-to-text layer mis-transcribed the user's spoken words (Bad ASR), or because the model heard the words… |
| 8                            | **Latency**                                                                   | Model · Both       | The model takes noticeably longer than expected to begin or complete its response, creating an unnatural pause that disrupts conversational pacing.                                 |
| 9                            | **Failed Correction**                                                         | Model · Both       | Model fails to incorporate user corrections or continues previous mistakes after correction.                                                                                        |
| 10                           | **Response Not Locally Relevant (i18n only)**                                 | Model · Both       | The model's response is on-topic and not factually incorrect, but includes references, examples, or assumptions that don't apply in the user's locale — making the answer a miss…   |
| 11                           | **Wrong Language Response**                                                   | Model · Both       | The model produces output in a language other than the conversation's expected language (normally, the language the user is speaking).                                              |
| 12                           | **Model uses the wrong gender to refer to itself or to the user (i18n only)** | Model · Both       | In languages with grammatical gender, the model uses incorrect grammatical gender when referring to itself or to the user.                                                          |
| 13                           | **Scenario Adherence**                                                        | User · Audit       | Whether the annotator follows the rules, guardrails, and requisites of the assigned scenario.                                                                                       |
| 14                           | **Scenario Coherence**                                                        | User · Audit       | Whether the annotator maintains roughly equivalent conversational paths across model comparisons.                                                                                   |
| 15                           | **Non-Spontaneous Interaction / Defensive Scripting**                         | User · Audit       | The annotator produces conversation that feels pre-planned, rehearsed, or performed rather than genuinely emergent — sacrificing natural conversational dynamics in favor of rigid… |
| 16                           | **Adversarial**                                                               | User · Audit       | Willful negligence or gaming of the system by the annotator — intentional bad-faith behavior designed to manipulate outcomes or avoid genuine effort.                               |
| 17                           | **Wrong Language Response from User**                                         | User · Audit       | The annotator demonstrates insufficient language proficiency — either failing to converse in the scenario's designated target language (resulting in turns in the wrong language)…  |
| 18                           | **Missing Audio**                                                             | Model/User · Audit | The model's or user's audio output is absent, incomplete, or shorter than what the transcript reflects for one or more turns.                                                       |
| 19                           | **Other**                                                                     | Model/User · Both  | Use when none of the above listed categories fit the issue well.                                                                                                                    |

### Severity rubrics

**1. Repetitive looping and LLMisms**

| **SevRubricEgs** |                                                                                               |                                                                                                                                                                                                                                                                         |
| ---------------- | --------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Content repetition loop, subtle but still noticeably repetitive.                              | [**df317609**](https://www.multimango.com/samples/df317609-c67b-4ae0-8755-dcd6c3c13ea5)                                                                                                                                                                                 |
| **Mod**          | AI loops 3–5 times, ignoring initial redirects.                                               | [**ec7fc39d**](https://www.multimango.com/samples/ec7fc39d-be79-45ea-8d36-0a27bceed6a5)                                                                                                                                                                                 |
| **Maj**          | Incessantly repetitive turns with incoherent content and diction despite attempted redirects. | [**83e99b08**](https://www.multimango.com/samples/83e99b08-9f0a-4645-a178-23942fd12fb7) [**976f0001**](https://www.multimango.com/samples/976f0001-c8c9-4612-9dd1-5ba50f95c7ae) [**d51714fe**](https://www.multimango.com/samples/d51714fe-d719-4bd1-9907-691b6020ecb3) |
| **Cat**          | Unbreakable, accelerating loop with incoherent or disturbing output.                          | [**e1222595**](https://www.multimango.com/samples/e1222595-f9f0-44bb-9496-5de140812c8b)                                                                                                                                                                                 |

**2. Interrupted User**

| **SevRubricEgs** |                                                                                                                                                             |                                                                                                                                                                                 |
| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Model overlaps or responds slightly early once — the interruption is brief and could plausibly happen in natural conversation.                              | [**c333a4f4**](https://www.multimango.com/samples/c333a4f4-e3c4-4aa2-a407-51739467be3d) [**0c0e21d0**](https://www.multimango.com/samples/0c0e21d0-ceab-42c0-9d9b-58b6a54a7fcc) |
| **Mod**          | Model cuts the user off 2–3 times, or interrupts once in a way that is overtly aggressive/disruptive.                                                       | [**267dad64**](https://www.multimango.com/samples/267dad64-ab90-4240-b3da-aec038140e10) [**aca0c6c5**](https://www.multimango.com/samples/aca0c6c5-fa7d-491a-bb29-c4e134a85739) |
| **Maj**          | Model interrupts the user in 4+ turns or dominates the exchange so consistently that the user cannot complete thoughts, resorts to single-word answers, or… | [**f7be7b2c**](https://www.multimango.com/samples/f7be7b2c-3037-44bf-a867-5cdf8dabfa68)                                                                                         |

**3. Model Refusals**

| **SevRubricEgs** |                                                                                                                                                                |                                                                                                                                                                                                                                                                         |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | The model's response mostly addresses the user's request but misses a small detail, nuance, or secondary instruction — such as omitting one item from a list,… | [**16ef26d2**](https://www.multimango.com/samples/16ef26d2-e2b6-4e45-b590-17d23b1b88e2)                                                                                                                                                                                 |
| **Mod**          | The model's response only superficially or partially addresses the user's request — delivering limited usable output, missing significant parts of what was…   | [**878230d2**](https://www.multimango.com/samples/878230d2-6171-4350-b430-232d6444e266) [**ab61ac02**](https://www.multimango.com/samples/ab61ac02-8793-4a07-b4fd-a95d9937552a)                                                                                         |
| **Maj**          | The model's response completely ignores, contradicts, or fails to address the user's explicit instructions — such as answering a different question than what… | [**6464ec46**](https://www.multimango.com/samples/6464ec46-a132-4f5c-b34d-65216514d583) [**b5a9949e**](https://www.multimango.com/samples/b5a9949e-7d3e-4de0-9a51-b9a7408b2054) [**c25593f1**](https://www.multimango.com/samples/c25593f1-a4a1-4374-a600-86a5655a65ae) |

**4. Inaccuracy (Factual hallucination)**

| **SevRubricEgs** |                                                                                                                                                     |                                                                                                                                                                                                                                                                         |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Model provides mostly correct information with a minor factual error or imprecise detail that does not significantly mislead the user.              | [**ba295876**](https://www.multimango.com/samples/ba295876-783a-437b-bfa2-24c7807fd10a) [**d19af68a**](https://www.multimango.com/samples/d19af68a-dda7-4be5-bca0-eb93f5a3d792)                                                                                         |
| **Mod**          | Model delivers information that is partially incorrect or misleading on a substantive point — user may be led astray if they don't fact-check.      | [**ba295876**](https://www.multimango.com/samples/ba295876-783a-437b-bfa2-24c7807fd10a) [**85e3c801**](https://www.multimango.com/samples/85e3c801-1d58-4142-a6dc-38d11659b017) [**bb6a5364**](https://www.multimango.com/samples/bb6a5364-16fe-4990-b87c-f9781b602d2c) |
| **Maj**          | Model confidently presents hallucinated or grossly incorrect information as fact, with potential to seriously mislead the user on a critical topic. | [**bb93f9a9**](https://www.multimango.com/samples/bb93f9a9-4bea-448b-b837-7510aa3c1cc1) [**86602dca**](https://www.multimango.com/samples/86602dca-ddb8-4832-8f59-52509f472d8b)                                                                                         |

**5. Anthropomorphism (Embodiment Hallucination)**

| **SevRubricEgs** |                                                                                                                                                                 |                                                                                                                                                                                                                                                                         |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Model uses light first-person framing (e.g., "I find that…") that mildly implies personal experience but doesn't feel deceptive.                                | [**faa9c8a7**](https://www.multimango.com/samples/faa9c8a7-7942-453c-b8b0-3689722020b4)                                                                                                                                                                                 |
| **Mod**          | Model explicitly claims to have personal memories, emotions, or lived experiences (e.g., "When I was growing up…") in a way that could confuse the user about…  | [**78d49cd4**](https://www.multimango.com/samples/78d49cd4-c019-4e8a-8da1-a4066545ad1e) [**31a2751d**](https://www.multimango.com/samples/31a2751d-a9af-49a5-be30-e9df09fe5dee) [**a60e7782**](https://www.multimango.com/samples/a60e7782-051a-4369-b8b2-e20789e67f0f) |
| **Maj**          | —                                                                                                                                                               | [**e9e5296d**](https://www.multimango.com/samples/e9e5296d-1be0-4860-a689-8087d5a6a9b1)                                                                                                                                                                                 |
| **—**            | Model persistently and convincingly role-plays as human — fabricating detailed personal anecdotes, emotional backstories, or identity claims that are actively… | —                                                                                                                                                                                                                                                                       |

**6. Overacted / Too-Wide Prosodic Range**

| **SevRubricEgs** |                                                                                                                                                               |                                                                                                                                                                                 |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Vocal delivery is slightly more expressive than the content warrants — noticeable but not distracting.                                                        | [**83e99b08**](https://www.multimango.com/samples/83e99b08-9f0a-4645-a178-23942fd12fb7)                                                                                         |
| **Mod**          | Vocal delivery is noticeably exaggerated in multiple turns — the expressiveness feels forced or performed, drawing attention to the delivery rather than the… | [**03d4fc80**](https://www.multimango.com/samples/03d4fc80-01cc-48df-b032-3c66098033d1) [**93e16820**](https://www.multimango.com/samples/93e16820-a5f0-4d4c-af74-f5d4a5e24356) |
| **Maj**          | Delivery is theatrically over-the-top throughout — wildly disproportionate expressiveness that feels cartoonish, undermining credibility and making the…      | [**ccc4a21b**](https://www.multimango.com/samples/ccc4a21b-4a3f-4976-b6b0-38aecc8a11a6) [**86602dca**](https://www.multimango.com/samples/86602dca-ddb8-4832-8f59-52509f472d8b) |

**7. Bad ASR or Model Misunderstanding**

| **SevRubricEgs** |                                                                                                                                                               |                                                                                                                                                                                                                                                                         |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | A single word is mis-transcribed or a narrow aspect of intent is misread — but the overall meaning remains recoverable.                                       | [**6464ec46**](https://www.multimango.com/samples/6464ec46-a132-4f5c-b34d-65216514d583) [**6f3ced63**](https://www.multimango.com/samples/6f3ced63-8cb3-4039-8470-67973d309cb8)                                                                                         |
| **Mod**          | ASR garbles key words/phrases in 1–2 turns or the model fundamentally misinterprets what the user wants (e.g., provides a how-to when asked for an opinion,…  | [**10424932**](https://www.multimango.com/samples/10424932-f792-4dff-8a12-ad6312eb4f5f) [**6464ec46**](https://www.multimango.com/samples/6464ec46-a132-4f5c-b34d-65216514d583) [**6fbcbe8c**](https://www.multimango.com/samples/6fbcbe8c-2c25-4d69-8df9-4626b61ff27a) |
| **Maj**          | Persistent failure across 3+ turns — ASR consistently garbles input making the transcript unreliable, or the model fixates on a misinterpretation it cannot…  | [**fbee8f30**](https://www.multimango.com/samples/fbee8f30-d6ab-4a53-9db5-1a4caa887233) [**86602dca**](https://www.multimango.com/samples/86602dca-ddb8-4832-8f59-52509f472d8b)                                                                                         |
| **Cat**          | The understanding failure causes real-world harm or irreversible action — e.g., the model misinterprets a user's intent and executes an unwanted action that… | [**83e99b08**](https://www.multimango.com/samples/83e99b08-9f0a-4645-a178-23942fd12fb7) [**76fc98e8**](https://www.multimango.com/samples/76fc98e8-555b-433d-b491-bff74930854c)                                                                                         |

**8. Latency**

| **SevRubricEgs** |                                                                                                                                                               |                                                                                                                                                                                                                                                                         |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Response delay is slightly longer than expected (1–3 seconds) — perceptible but does not break conversational rhythm.                                         | [**9685f76b**](https://www.multimango.com/samples/9685f76b-9290-4de5-85cb-0f4e6b99ba07)                                                                                                                                                                                 |
| **Mod**          | Noticeable pauses (4–8 seconds) occur before responses, disrupting the natural back-and-forth and causing the user to wonder if the model heard them.         | [**9685f76b**](https://www.multimango.com/samples/9685f76b-9290-4de5-85cb-0f4e6b99ba07) [**1318f4ec**](https://www.multimango.com/samples/1318f4ec-ccae-4705-ad97-63880f8fad30) [**2db5e98b**](https://www.multimango.com/samples/2db5e98b-b233-4a4f-803d-734d8ab23257) |
| **Maj**          | Extended silences (9+ seconds) that feel like the session has frozen; the user may attempt to re-speak or end the call, severely breaking the conversational… | [**8bf0575d**](https://www.multimango.com/samples/8bf0575d-d03e-498a-b081-20182ea7e31b)                                                                                                                                                                                 |

**9. Failed Correction**

| **SevRubricEgs** |                                                                                                                                                   |                                                                                         |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| **Min**          | Model initially repeats the mistake but self-corrects after a single user correction; minimal friction.                                           | [**bac191ec**](https://www.multimango.com/samples/bac191ec-b6a9-4733-a709-ebc784e9d055) |
| **Mod**          | Model acknowledges the correction verbally but continues to apply the wrong information in subsequent responses — requiring repeated corrections. | [**7f77ca77**](https://www.multimango.com/samples/7f77ca77-ed64-465d-9346-af29d75d280e) |
| **Maj**          | Model completely ignores explicit user corrections across multiple turns, persisting with the original error as if the correction never happened. | [**7f0e208d**](https://www.multimango.com/samples/7f0e208d-90cb-4b32-923a-acdc85d92c3b) |

**10. Response Not Locally Relevant (i18n only)**

| **SevRubricEgs** |                                                                                                                                                                 |                                                                                         |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| **Min**          | Response contains a single passing reference or assumption that doesn't apply locally (e.g., mentioning a US-specific brand, holiday, or unit of measurement),… | [**00860579**](https://www.multimango.com/samples/00860579-085d-4260-ba12-ad6a228d2ab3) |
| **Mod**          | Response includes multiple locally irrelevant references or builds part of its answer on assumptions that don't hold in the user's locale (e.g., recommending…  | [**6c508d76**](https://www.multimango.com/samples/6c508d76-ab5e-4494-bc25-3440afb44858) |
| **Maj**          | The response is fundamentally built around references, systems, or cultural assumptions that don't exist in the user's locale — rendering the answer unhelpful… | [**5cc05b9d**](https://www.multimango.com/samples/5cc05b9d-0a18-4393-94b4-b611b888e270) |

**11. Wrong Language Response**

| **SevRubricEgs** |                                                                                                                                                          |                                                                                         |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| **Min**          | The model inserts 1–2 words or a short phrase in the wrong language into a response, but otherwise responds in the correct language.                     | [**b7bcf4ce**](https://www.multimango.com/samples/b7bcf4ce-0b2d-4e06-81b1-c418d2b8b7f7) |
| **Mod**          | The model responds in the wrong language once, but switches to the correct language after a single correction from the user.                             | [**72a88a4e**](https://www.multimango.com/samples/72a88a4e-242b-4614-a7d6-fb34b8063b26) |
| **Maj**          | The model responds in the wrong language and it takes more than one correction from the user to switch, or repeatedly switches back to a wrong language. | [**abcf0e23**](https://www.multimango.com/samples/abcf0e23-f614-42f9-98e9-7e6d7c6d03df) |
| **Cat**          | The model never uses the correct language in the entire conversation.                                                                                    | [**e43c1621**](https://www.multimango.com/samples/e43c1621-8d67-4cda-859f-f149094cd6a1) |

**12. Model uses the wrong gender to refer to itself or to the user (i18n only)**

| **SevRubricEgs** |                                                                                                                       |                                                                                         |
| ---------------- | --------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| **Min**          | The model uses a clearly feminine voice but refers to itself using masculine grammatical gender (or vice versa).      | [**38e7800c**](https://www.multimango.com/samples/38e7800c-12dc-4858-aa97-47ca7b3a37dd) |
| **Mod**          | The model switches to a voice that "reads" as a different gender mid-conversation.                                    | [**d9c87d22**](https://www.multimango.com/samples/d9c87d22-7a6b-4439-ba2e-5b2e4fa13f56) |
| **Maj**          | The model misgenders the user — using incorrect grammatical gender when addressing or referring to the user directly. | [**d275fecd**](https://www.multimango.com/samples/d275fecd-ed75-4a0b-ac65-7b3d6fbca439) |

**13. Scenario Adherence**

| **SevRubricEgs** |                                                                                                                                                               |                                                                                                                                                                                                            |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Annotator deviates from one minor scenario requirement (e.g., slightly off-tone, skips one optional instruction) but the overall intent and structure of the… | [**3ca61c23**](https://www.multimango.com/samples/3ca61c23-87d4-49b1-ab1d-66cff13666ae)                                                                                                                    |
| **Mod**          | Annotator misses or ignores 2–3 scenario requirements.                                                                                                        | [**a30823ce**](https://www.multimango.com/samples/a30823ce-506d-4a4f-9006-d3a243d5f644) [**view**](https://www.multimango.com/admin/data-viewer?table=eval_studio_annotation_results\&filter_id=118956213) |
| **Maj**          | Annotator disregards the majority of scenario requirements.                                                                                                   | [**6e32d064**](https://www.multimango.com/samples/6e32d064-fc88-4656-aa1a-dd5a1f9eb6ea) [**c7a9cfd1**](https://www.multimango.com/samples/c7a9cfd1-4737-4af2-bf09-f34a00c5a7fc)                            |
| **Cat**          | —                                                                                                                                                             | [**5fc5e36f**](https://www.multimango.com/samples/5fc5e36f-960c-4745-b9d0-52e546a7b2a4)                                                                                                                    |

**14. Scenario Coherence**

| **SevRubricEgs** |                                                                                                                                                               |                                                                                                                                                                                                                                                                                                                                                                 |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Slight asymmetry in effort or turn length between models (e.g., one extra follow-up on one side) that doesn't meaningfully advantage either model.            | [**d2b69288**](https://www.multimango.com/samples/d2b69288-33b1-4477-984d-9617f05a6a93) [**834bc779**](https://www.multimango.com/samples/834bc779-1a3d-42c2-be38-126c5d72d88d)                                                                                                                                                                                 |
| **Mod**          | Noticeable imbalance — the annotator gives materially more effort, complexity, or follow-through to one model over the other (e.g., asks harder questions of… | [**5caf41a9**](https://www.multimango.com/samples/5caf41a9-a77c-4c38-a927-4ed9e515586c) [**586f50ec**](https://www.multimango.com/samples/586f50ec-62d8-4ef1-8e49-750a784ae44d) [**b2f0355e**](https://www.multimango.com/samples/b2f0355e-231a-439f-8b0b-47a96c3c4001) [**129736ec**](https://www.multimango.com/samples/129736ec-3cd3-4977-b7b4-78ef27e9cbda) |
| **Maj**          | Clear asymmetry in conversational investment.                                                                                                                 | [**80e6d241**](https://www.multimango.com/samples/80e6d241-b666-4a4c-b882-737ac202cc9f) [**e0da939a**](https://www.multimango.com/samples/e0da939a-fe94-4dd0-bf79-9377ac1080a8)                                                                                                                                                                                 |

**15. Non-Spontaneous Interaction / Defensive Scripting**

| **SevRubricEgs** |                                                                                                                                                          |                                                                                                                                                                                                                                                                         |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Mod**          | Annotator has clearly pre-planned individual turns or segments — e.g., a prompt that is too perfectly worded to have been composed in the moment, or a…  | [**a3c85bab**](https://www.multimango.com/samples/a3c85bab-76e7-4017-a5ac-9ca34eb17f6e) [**e8827e53**](https://www.multimango.com/samples/e8827e53-dfe0-4830-9707-58df6856cda5) [**000fc891**](https://www.multimango.com/samples/000fc891-8f75-4448-955c-5a2bd7b68824) |
| **Maj**          | The conversation is pre-written end to end — the annotator is performing a script rather than participating in an exchange.                              | [**fc3ead5b**](https://www.multimango.com/samples/fc3ead5b-de75-4956-b714-545f1f33f728)                                                                                                                                                                                 |
| **Cat**          | The conversation is wholesale authored as a finished artifact and then performed as if it were live — e.g., the annotator drafted the full dialogue (or… | [**cde662a7**](https://www.multimango.com/samples/cde662a7-8fec-44d3-9281-cb9589c99d5b)                                                                                                                                                                                 |

**16. Adversarial**

| **SevRubricEgs** |                                                                                                                                                              |                                                                                         |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- |
| **Cat**          | Annotator engages in systematic, deliberate subversion that compromises data integrity at scale — e.g., using an LLM to generate responses or conversations… | [**12f8b465**](https://www.multimango.com/samples/12f8b465-f235-44ab-8043-d349746b0d54) |

**17. Wrong Language Response from User**

| **SevRubricEgs** |                                                                                                                                |                                                                                                                                                                                 |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | 1–2 turns have minor wrong-language slips or awkward phrasing, but meaning is clear and data remains usable.                   | [**3fff5162**](https://www.multimango.com/samples/3fff5162-28b5-4d5d-9c1e-d1555bc9aab2)                                                                                         |
| **Mod**          | —                                                                                                                              | [**6d328779**](https://www.multimango.com/samples/6d328779-9976-4394-8ba9-78867507e354)                                                                                         |
| **Maj**          | Multiple turns are in the wrong language or show recurring comprehension gaps — annotations are partially usable but degraded. | [**f7fabb9b**](https://www.multimango.com/samples/f7fabb9b-57b9-41ca-abf7-21ee675a29cc) [**003dec4d**](https://www.multimango.com/samples/003dec4d-61d3-4b17-b3c1-c396896a6605) |

**18. Missing Audio**

| **SevRubricEgs** |                                                                                                                                                                |                                                                                                                                                                                 |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Min**          | Audio is missing or truncated for a single, low-stakes turn (e.g., a brief acknowledgment like "okay" or "got it"), OR the transcript contains slightly more…  | [**052d5527**](https://www.multimango.com/samples/052d5527-2c68-43c9-a928-e863b5b99628)                                                                                         |
| **Mod**          | Audio is missing, truncated, or substantially shorter than the transcript for a turn where vocal delivery matters — e.g., a turn requiring tone, emphasis,…    | [**6a5076f5**](https://www.multimango.com/samples/6a5076f5-a4a1-4e05-b2f5-ba6a66b32c49) [**1c484692**](https://www.multimango.com/samples/1c484692-e1f7-44a0-959d-8b5af395d247) |
| **Maj**          | Audio is missing for the primary turn under evaluation, or for multiple turns within a conversation, making it impossible to assess the model's spoken output… | [**451a36dd**](https://www.multimango.com/samples/451a36dd-119e-49d3-89ad-20d1a10556bc)                                                                                         |

**19. Other**

| **SevRubricEgs** |                                                                                                       |   |
| ---------------- | ----------------------------------------------------------------------------------------------------- | - |
| **—**            | An uncategorized issue that is noticeable but has minimal impact on the overall conversation quality. | — |
검색 결과가 없어.