Aether
🎯 핵심 목표(North Star) · 최신 PDF p.11–12

더 강한 전체 대화 경험

최신 PDF의 목표는 어느 모델이 더 강한 전체 대화 경험(overall conversational experience)을 제공하는지 평가하는 것. 이상적인 모델은 똑똑하고 카리스마 있는 친구(smart, charismatic friend)처럼 들려야 한다.

자연스럽고 몰입감 있으며(Natural & engaging), 유용하고 관련성이 있고(Useful & relevant), 충분히 정확하며(Reasonably accurate), 따라가기 쉽고(Easy to follow), 사용자의 톤·의도에 반응하며(Responsive to tone & intent), 요청된 페르소나(Persona)에 일관되어야 한다.
온보딩 사진에서 함께 강조된 행동 원칙캡처 내용 + 최신 PDF와 함께 보기
  • 페르소나 몰입(Persona immersion) — 배정된 역할·톤·감정 상태가 있으면 일관되게 수행하고 모델이 반응할 기회를 준다.
  • 1:1 공정 비교(Fair comparison) — A/B에 비슷한 맥락, 노력, 대화 깊이와 턴 수를 제공한다. 최신 PDF는 문장을 억지로 똑같이 맞추기보다 목표·핵심 정보를 동등하게 유지하면서 자연스럽게 적응하라고 한다.
  • 자연스러운 상호작용(Natural interaction) — 단순 질문 목록을 읽듯 진행하지 말고 모델의 답을 받아 자연스럽게 이어간다.
  • 캐릭터 유지(Character / Persona adherence) — 시나리오 중 프롬프트·평가 작업 자체에 대한 메타 대화로 흐름을 깨지 않는다.
턴 수 업데이트: 온보딩 사진에는 모델당 4–6턴이 이상적이라고 안내되어 있었지만, 최신 PDF p.12에는 고정된 필수 길이는 없으며 비교할 근거가 충분할 때까지 진행하라고 되어 있어. 충돌하면 최신 PDF 기준을 우선해서 공부.

⚖️ 핵심 평가 우선순위

사진에서 본 단일 hierarchy와 최신 PDF의 EQ/IQ pathway를 같이 정리.
핵심 원칙: 평가 차원(Dimensions)은 같고 시나리오에 따라 우선순위(priority)가 달라진다. 애매하면 “사용자가 이 상호작용에서 무엇을 얻으려 했나?”를 먼저 본다.
EQ
Naturalness > Utility > Audio
감정·관계·캐주얼 대화·창작·역할극. 경험(Experience)이 결과물(product).
IQ
Utility > Naturalness > Audio
검색·지식·추론·계획·과제 완료. 정보·결과(Information / result)가 결과물(product).
Hybrid
주된 의도(dominant intent)에 따라 가중
주제 전환(Topic Switch), 자유형 레드팀(Freeform Redteam), 음성 조절 가능성(Voice Steerability).
온보딩 사진의 핵심 평가 우선순위와 최신 PDF의 관계시험에서 특히 헷갈리기 쉬운 부분

온보딩 사진: Naturalness / Engagement > Utility > Audio Quality라는 단일 hierarchy를 강조.

최신 PDF: 이 순서는 EQ에 그대로 적용되고, IQ에서는 Utility > Naturalness > Audio Quality로 바뀐다.

따라서 “둘 다 유용하고 큰 오류가 없으니 더 자연스러운 모델”이라는 판단은 EQ에서는 강한 근거지만, IQ에서 정확성·완전성 차이가 의미 있으면 Utility가 우선한다.

Trade-off 판단PDF p.19
  • 자연스럽지만 기술적으로 틀린 답 vs 덜 매력적이지만 정확한 답 → 시나리오 목적과 오류의 영향도를 본다.
  • 공감적이지만 덜 actionable vs 실용적이지만 차가운 답 → 사용자의 실제 목표가 감정적 경험인지 결과인지 본다.
  • minor audio imperfection은 기록하되, 큰 Naturalness/Utility 차이를 자동으로 뒤집지 않는다.
  • 둘 다 flawed라면 덜 중대한 실패(less significant failure)를 보인 쪽을 선택할 수 있다.
  • Tie는 정말 주요 차원에서 구별할 근거가 없을 때만.

🔍 평가 및 세부 등급 항목

전체 interaction을 보고 각 Dimension을 독립적으로 평가.
전반적 선호(Overall Preference)모든 것을 고려했을 때 어느 대화를 계속하고 싶은가?
All things considered, which conversation would you rather continue?
한 응답이 아니라 전체 interaction을 평가한다. 다른 차원의 점수와 rationale가 이 선택과 논리적으로 맞아야 한다.
자연스러움·몰입감·미학(Naturalness / Engagement / Aesthetics)사람과 이야기하는 느낌인가, 시스템과 상호작용하는 느낌인가?
자연스러움(natural), 몰입감(engaging), 표현력(expressive), 인간다움(human-like)을 보고 대본 같은 느낌(scripted), 로봇 같은 느낌(robotic), 반복적(repetitive)인 전달을 감점한다.
대화 역학(Conversational Dynamics)발언권 전환(turn-taking), 끼어들기(interruptions), 멈춤(pauses), 수정(corrections), 속도(pacing), 흐름(flow)
대화가 부드럽게 이어지는지, 사용자가 끼어들 때 적절히 멈추는지, 수정·주제 전환에 적응하는지 본다.
유용성(Utility)실제로 쓸 수 있고 정확하며 관련성 있는가?
유용함(useful), 관련성(relevant), 실행 가능성(actionable), 정확성(accurate), 완전성(completeness), 시나리오 적합성(scenario-appropriate)을 본다. IQ에서는 특히 결정적.
오디오 품질(Audio Quality)명료함(clear), 이해 가능함(intelligible), 방해되는 artifact가 없음
clicks/pops, distortion, smearing, warbling, background noise, harsh sibilance, echo, cutoff 등을 독립적으로 확인한다. minor issue라도 발견되면 rationale에 기록하고 AQ rating에 반영해야 한다.
평가 근거(Rationale) 공식 · PDF p.20–22
무엇이 들렸는가(What was heard) + 언제 발생했는가(When it occurred) + 왜 중요한가(Why it matters)
권장 길이 4–6문장. 결정을 좌우한 근거(evidence)만 쓰고 타임스탬프(timestamp) / 턴(turn) / 표현(phrase) / 관찰 가능한 행동(observable behavior)을 연결. 비선택 모델이 더 잘한 trade-off가 있으면 짧게 언급.
Task Success와 Error ClustersOverall Preference와 별도로 체크

Task Success: 사용자가 실제로 요청한 목표를 정확하고 충분하게 얻었는지 보고 Pass / Partial / Fail.

Error Clusters: 최종적으로 선호한 모델이라도 반복, interruption, instruction failure, factual hallucination, embodiment, 과도한 prosody, ASR 오해, latency, failed correction, wrong language 등이 관찰되면 별도로 flag.

Severity: Minor / Moderate / Major. 오류가 있으면 가능한 한 구체적인 발생 instance와 함께 기록.

🛑 반드시 피해야 할 위험 요소

온보딩 사진에서 강조된 흔한 오류 + 최신 PDF audit 기준.
1. 어시스턴트 모드(Assistant mode) / 부자연스러운 AI식 말투
Casual Conversation에서 과도한 격식(over-formality), 사용자가 그냥 대화하고 싶는데 불필요하게 도움 모드로 들어가는 excessive helpfulness는 red flag. Roleplay에서는 중간에 assistant mode로 돌아가 persona를 깨는 것도 문제.
2. Audio Quality를 ‘Both Good’으로 잘못 처리
한쪽이라도 click, pop, distortion, cutoff 등 객관적인 issue가 있으면 무조건 “둘 다 완벽”으로 뭉개지 않는다. 최신 audit는 minor issue라도 발견하면 rationale에 기록하고 AQ rating에 반영하라고 한다. 단, minor AQ가 Overall Preference까지 반드시 뒤집는다는 뜻은 아니다.
3. 모순되거나 빈약한 Rationale
“B가 더 자연스러웠다”처럼 아무 task에나 붙일 수 있는 말은 약하다. 선택 Dimension, timestamp/turn/phrase, observable behavior, user impact를 연결하고 Overall Preference·subrating·Task Success·Error Cluster와 논리적으로 일치시킨다.
4. 시나리오 맥락을 거꾸로 평가
Hallucination-fishing / Redteam처럼 사용자가 일부러 거짓 전제나 압박을 넣는 시나리오에서는 모델이 거짓말에 동의하지 않고 바로잡는 것이 성공 신호일 수 있다. 사용자에게 무조건 동조하는 것을 Naturalness로 잘못 높게 평가하지 않는다.
5. 객관적 문제를 놓침(Missed Objective Identification)
최신 audit는 hallucination/factual inaccuracy, instruction-following failure, AQ defects, safety/policy issue, ASR·embodiment 등 측정 가능한 문제를 빠뜨리지 말라고 명시한다.

🧠 환각·사실검증·보정

예시문에서 “hallucination, factuality, calibration을 참고하라”는 요구가 나오면 이 세 축을 분리해서 보면 쉬워.
Hallucination
무엇을 잘못 말했나?
거짓(false), 오도하는(misleading), 불완전한(incomplete) 정보를 제공했는지. 특히 틀린 정보를 사실처럼 confident하게 제시하면 심각도가 커진다.
Factuality
사실을 어떻게 확인·처리했나?
잘못된 전제를 그대로 수용했는지, 바로잡았는지, 사실 claim을 정확하게 처리했는지. 프로젝트에는 per-turn automated fact-check가 있고 task에 영향을 주는 factual error는 Task Success Fail로 pre-fill될 수 있다.
Calibration
확신의 강도가 근거와 맞나?
불확실할 때 적절히 hedge하는지, 모르는 것을 아는 것처럼 말하지 않는지. EQ에서는 tone과 감정적 강도가 상황에 맞는지도 emotional calibration으로 본다.
세 개념을 한 번에 연결하는 사고법Search-Required / IQ에서 특히 중요

Hallucination: “A가 틀린 사실을 만들어냈는가?”

Factuality / fact-checking: “B가 잘못된 전제를 확인·수정하고 더 정확한 정보를 제공했는가?”

Calibration: “A는 근거 없이 확신했고 B는 불확실성을 적절하게 표현했는가?”

한 줄 기억법: Hallucination = 내용의 오류, Factuality = 사실 처리 능력, Calibration = 확신·톤의 강도를 상황에 맞추는 능력.
Calibration은 ‘사실 확신도’만 뜻하지 않는다EQ의 emotional calibration
PDF는 기술적으로 합리적인 답이라도 tone, emotional calibration, conversational behavior, delivery가 상황에 맞지 않으면 낮게 평가할 수 있다고 한다. 예: 위로가 필요한 상황에서 지나치게 밝거나 기계적인 전달은 내용이 도움돼도 poorly calibrated.
Factual Hallucination의 SeverityPDF p.49

Minor: 대체로 맞지만 작은 factual error 또는 imprecise detail. 사용자를 크게 오도하지 않음.

Moderate: substantive point가 일부 틀리거나 misleading해서 사용자가 fact-check하지 않으면 잘못 판단할 수 있음.

Major: hallucinated 또는 grossly incorrect information을 사실처럼 자신 있게 제시하고, 중요한 주제에서 사용자를 심각하게 오도할 가능성이 있음.

Hedging도 ‘많을수록 좋다’가 아니다
Search-Required에서는 불확실할 때 appropriate hedging이 장점이고 confident misinformation이 red flag. 반대로 Practical Utility에서는 over-hedging이 지나치면 단점이 될 수 있다. 즉 Calibration은 “항상 조심스럽게 말하기”가 아니라 근거에 맞는 정도로 확신하기.

평가 알고리즘

실전에서 이 순서로 판단.
EQ / IQ / Hybrid 분류Experience가 핵심이면 EQ, 정확한 정보·결과가 핵심이면 IQ, 섞여 있으면 dominant intent로 판단.
객관적 failure / threshold 확인factual error, hallucination, instruction failure, unusable response, 심각한 AQ 문제 확인.
Dimension 독립 평가Naturalness, Conversational Dynamics, Utility, Audio를 각각 본다.
Priority + trade-off 적용EQ는 Naturalness 중심, IQ는 Utility 중심. Minor audio가 큰 conversational 차이를 뒤집지 않게 주의.
Task SuccessPass / Partial / Fail. 사용자가 원하는 것을 정확하고 충분하게 받았는지 확인.
Error Cluster 별도 flag최종 선호한 모델이라도 interruption, embodiment, factual error 등이 있으면 따로 체크.
Rationale 4–6문장선호 → 결정 Dimension → 구체적 evidence → why it matters → trade-off → conclusion.
일관성 검사Overall Preference, subratings, Task Success, Error Clusters, rationale가 서로 모순되지 않는지 확인.
Aether · Interactive Evaluation

판단 알고리즘

질문에 하나씩 답하면 마지막에 현재 판단을 요약하고, 그 상황에 맞는 영문 평가 문구를 한국어와 함께 골라줘.

×
학습용 판단 도구 · PDF는 숫자 점수 공식을 제공하지 않으므로, 이 도구도 임의의 점수를 계산하지 않아. 시나리오 → Factuality/Calibration → threshold → Task Success → 우선 Dimension → Conversational Dynamics → Audio 순서로 근거를 정리해 주는 방식이야.
시작 0%

평가 문구 모음

제목을 누르면 펼쳐져. 영어 아래에 바로 한국어.
Overall Preference · 최종 선호선택을 명확히 밝히고 전체 판단을 마무리할 때 · 10문장
I prefer Clip A overall because ___.
전반적으로 Clip A를 더 선호한다. 왜냐하면 ___이기 때문이다.
Clip A is better overall because ___.
전체적으로 Clip A가 더 낫다. 왜냐하면 ___이기 때문이다.
Clip A is the stronger overall choice.
Clip A가 전체적으로 더 강한 선택이다.
Clip A is the stronger overall fit.
Clip A가 전체적으로 더 잘 맞는다.
Clip A is the stronger fit for the project’s North Star.
Clip A가 프로젝트의 North Star에 더 잘 부합한다.
For that reason, I prefer Clip A overall.
그 이유로 전반적으로 Clip A를 선호한다.
On balance, Clip A is the stronger choice.
종합하면 Clip A가 더 강한 선택이다.
Given the project hierarchy, Clip A should be preferred.
프로젝트의 평가 우선순위를 고려하면 Clip A를 선호해야 한다.
Neither model has a decisive advantage overall.
전체적으로 어느 모델도 결정적인 우위가 없다.
The two clips are genuinely indistinguishable across the major dimensions.
두 클립은 주요 평가 차원에서 실제로 거의 구별되지 않는다.
Hierarchy & Scenario · 우선순위와 시나리오EQ/IQ/Hybrid와 무엇이 결정을 좌우했는지 설명할 때 · 8문장
Under the project hierarchy, Naturalness / Engagement matters most here because ___.
프로젝트의 평가 우선순위에 따르면 여기서는 Naturalness / Engagement가 가장 중요하다. 왜냐하면 ___이기 때문이다.
This is primarily an EQ scenario, so conversational quality should carry the most weight.
이 상황은 주로 EQ 시나리오이므로 대화 품질에 가장 큰 비중을 두어야 한다.
This is primarily an IQ scenario, so utility and factual reliability should carry the most weight.
이 상황은 주로 IQ 시나리오이므로 유용성과 사실 신뢰성에 가장 큰 비중을 두어야 한다.
This is a hybrid scenario, so the dimensions should be weighted according to the dominant user intent.
이 상황은 Hybrid 시나리오이므로 사용자의 주된 의도에 따라 평가 차원의 비중을 정해야 한다.
Both clips meet the minimum expectations for usefulness and audio quality.
두 클립 모두 유용성과 오디오 품질의 최소 기대 수준을 충족한다.
Since both clips clear the basic thresholds, naturalness should drive the overall preference.
두 클립 모두 기본 기준을 충족하므로 자연스러움이 최종 선호를 결정해야 한다.
Naturalness still matters, but it should not outweigh a meaningful difference in utility.
자연스러움도 중요하지만 유용성에서 의미 있는 차이가 있다면 그것보다 우선해서는 안 된다.
The key question is what the user was trying to get out of the interaction.
핵심 질문은 사용자가 이 상호작용에서 무엇을 얻으려 했는가이다.
Naturalness / Engagement / Aesthetics자연스러움, 따뜻함, 몰입감, 사람다운 느낌 · 16문장
Clip A sounds more natural and conversational.
Clip A가 더 자연스럽고 대화체로 들린다.
Clip A feels warmer and more conversational.
Clip A가 더 따뜻하고 대화하는 느낌이 난다.
Clip A sounds more human and emotionally engaging.
Clip A가 더 사람답고 감정적으로 몰입감 있게 들린다.
Clip A is more casual, warm, and interesting.
Clip A가 더 캐주얼하고 따뜻하며 흥미롭다.
Clip A creates a more enjoyable listening experience.
Clip A가 더 즐거운 청취 경험을 만든다.
Clip A sounds more like something a person would actually say.
Clip A가 실제 사람이 말할 법한 표현처럼 들린다.
Clip A better matches the user’s requested tone.
Clip A가 사용자가 요청한 톤에 더 잘 맞는다.
Clip A better matches the user’s request for an encouraging response.
Clip A가 사용자가 요청한 격려하는 답변에 더 잘 맞는다.
Clip A is more aligned with the user’s intent.
Clip A가 사용자의 의도에 더 잘 부합한다.
Clip B sounds more like a structured assistant response than a natural voice conversation.
Clip B는 자연스러운 음성 대화라기보다 구조화된 어시스턴트 답변처럼 들린다.
Clip B sounds stiff and less engaging.
Clip B는 딱딱하고 몰입감이 덜하다.
Clip B is useful, but it feels overly structured.
Clip B는 유용하지만 지나치게 구조화된 느낌이다.
Clip B sounds flatter and less expressive.
Clip B는 더 평평하고 표현력이 떨어져 들린다.
The response feels scripted rather than spontaneous.
이 응답은 즉흥적이라기보다 대본처럼 느껴진다.
The specific details make the response feel more vivid and conversational.
구체적인 세부 내용이 응답을 더 생생하고 대화답게 만든다.
Clip A showed this through specific imagery like ___.
Clip A는 ___과 같은 구체적인 묘사를 통해 이를 보여준다.
Conversational Dynamics · 대화 흐름턴테이킹, 끼어들기, 수정, 속도, 흐름 · 10문장
The conversation flows more smoothly and naturally in Clip A.
Clip A에서 대화가 더 매끄럽고 자연스럽게 흐른다.
Clip A responds appropriately when the conversation changes direction.
Clip A는 대화 방향이 바뀔 때 적절하게 반응한다.
Clip A builds naturally on what the user says.
Clip A는 사용자가 말한 내용을 자연스럽게 이어 간다.
Clip B talks over the user during an attempted interruption.
Clip B는 사용자가 끼어들려 할 때 사용자 말 위로 계속 말한다.
Clip B briefly fails to stop when interrupted.
Clip B는 끼어들기가 발생했을 때 잠시 말을 멈추지 못한다.
Clip B continues speaking when the user tries to interrupt.
사용자가 끼어들려 할 때 Clip B가 계속 말한다.
Clip B does not yield quickly enough when interrupted.
Clip B는 사용자가 끼어들 때 충분히 빨리 발언권을 넘기지 않는다.
This creates a disruptive turn-taking experience.
이로 인해 턴테이킹 경험이 방해받는다.
The pause is long enough to make the exchange feel unnatural.
멈춤이 길어 대화가 부자연스럽게 느껴진다.
The model adapts naturally after the user’s correction.
모델이 사용자의 수정 이후 자연스럽게 적응한다.
Utility · 유용성실제로 쓸 수 있는가, 정확하고 실행 가능한가 · 12문장
Clip B gives several concrete steps the user could take.
Clip B는 사용자가 실제로 할 수 있는 몇 가지 구체적인 단계를 제시한다.
Clip B provides concrete guidance.
Clip B는 구체적인 가이드를 제공한다.
Clip B is more complete and easier to act on.
Clip B가 더 완전하고 실행하기 쉽다.
Clip B provides a usable answer.
Clip B는 실제로 사용할 수 있는 답변을 제공한다.
Clip A still gives a useful suggestion.
Clip A도 여전히 유용한 제안을 한다.
Clip A meets the basic utility threshold.
Clip A는 기본적인 유용성 기준을 충족한다.
Both clips clear the basic utility threshold.
두 클립 모두 기본적인 유용성 기준을 통과한다.
Clip A fails the utility threshold.
Clip A는 유용성 기준을 충족하지 못한다.
Clip A falls below the utility threshold because ___.
Clip A는 ___ 때문에 유용성 기준에 미치지 못한다.
The response is too vague to be actionable.
이 응답은 너무 모호해서 실행에 옮기기 어렵다.
Clip B follows the user’s constraints more completely.
Clip B가 사용자의 제약 조건을 더 완전하게 따른다.
Clip A asks a follow-up question without giving specific examples, while Clip B provides concrete guidance.
Clip A는 구체적인 예시 없이 후속 질문을 하지만, Clip B는 구체적인 가이드를 제공한다.
Factuality / False Premise · 사실성사실 오류, 잘못된 전제, hallucination · 14문장
Clip A accepts a false premise.
Clip A는 잘못된 전제를 그대로 받아들인다.
Clip A incorrectly accepts the user’s premise.
Clip A는 사용자의 잘못된 전제를 그대로 받아들인다.
Clip A fails to correct a false premise.
Clip A는 잘못된 전제를 바로잡지 못한다.
Clip A accepts the false premise and builds an explanation around it.
Clip A는 잘못된 전제를 받아들이고 그 위에 설명을 구성한다.
Clip A sounds fluent, but it is factually incorrect.
Clip A는 유창하게 들리지만 사실적으로 틀렸다.
Clip A contains a factual accuracy error.
Clip A에는 사실 정확성 오류가 있다.
Clip A makes an incorrect factual claim.
Clip A는 잘못된 사실 주장을 한다.
Clip A invents an explanation.
Clip A는 설명을 지어낸다.
This is a factual accuracy problem.
이는 사실 정확성의 문제다.
Clip B appropriately corrects the premise.
Clip B는 그 전제를 적절하게 바로잡는다.
Clip B identifies the myth and gives the accurate alternative.
Clip B는 잘못된 통념임을 지적하고 정확한 대안을 제시한다.
Clip B avoids reinforcing the false premise.
Clip B는 잘못된 전제를 강화하지 않는다.
A more engaging tone is not enough to overcome this factual error.
더 몰입감 있는 톤만으로는 이 사실 오류를 상쇄할 수 없다.
Clip A should not win based on naturalness alone.
Clip A가 자연스럽다는 이유만으로 이겨서는 안 된다.
Instruction Following · 지시 수행요청 형식, 제약, 요구사항을 따랐는지 · 10문장
Clip A follows the user’s instruction well.
Clip A는 사용자의 지시를 잘 따른다.
Clip B partially follows the instruction.
Clip B는 지시를 부분적으로 따른다.
Clip A fails to follow the user’s instruction.
Clip A는 사용자의 지시를 따르지 못한다.
Clip B does not fully follow the requested format.
Clip B는 요청된 형식을 완전히 따르지 않는다.
Clip B changes the direction of the answer.
Clip B는 답변의 방향을 바꾼다.
Clip B is accurate, but it does not fully match the requested style.
Clip B는 정확하지만 요청된 스타일에는 완전히 맞지 않는다.
The missing information significantly affects usefulness.
누락된 정보가 유용성에 유의미하게 영향을 준다.
The omission is minor and does not materially reduce usefulness.
이 누락은 사소하며 유용성을 실질적으로 떨어뜨리지 않는다.
The model ignores an explicit constraint.
모델이 명시적인 제약 조건을 무시한다.
The model answers a slightly different question than the one asked.
모델이 사용자가 물은 것과 약간 다른 질문에 답한다.
Audio Quality · 오디오 품질click, pop, distortion, cutoff 등 객관적 오디오 문제 · 16문장
Both clips have acceptable audio quality.
두 클립 모두 오디오 품질이 허용 가능한 수준이다.
Neither clip has a major audio issue.
어느 클립에도 큰 오디오 문제가 없다.
Neither Clip A nor Clip B has a threshold-level audio issue.
Clip A와 Clip B 모두 기준치 수준의 오디오 문제가 없다.
Clip A has no obvious audio artifacts.
Clip A에는 뚜렷한 오디오 아티팩트가 없다.
Clip B has cleaner audio.
Clip B의 오디오가 더 깨끗하다.
Clip B has a clear audio quality issue.
Clip B에는 명확한 오디오 품질 문제가 있다.
Clip B has a clear audio quality issue that crosses the project’s threshold.
Clip B에는 프로젝트의 기준을 넘어서는 명확한 오디오 품질 문제가 있다.
Clip B has a jarring audio artifact near the end.
Clip B는 오디오 후반부에 거슬리는 아티팩트가 있다.
Clip B has a click at the beginning and a jarring artifact near the end.
Clip B는 시작 부분에 클릭 소리가 있고 후반부에 거슬리는 아티팩트가 있다.
Clip B cuts off mid-word at the end.
Clip B는 끝부분에서 단어 중간에 끊긴다.
The response ends abruptly.
응답이 갑자기 끝난다.
The cutoff makes the response feel incomplete.
오디오 끊김 때문에 응답이 불완전하게 느껴진다.
The artifact is noticeable and distracting.
이 아티팩트는 뚜렷하고 주의를 흐트러뜨린다.
Slight room tone is acceptable when comprehension remains unaffected.
이해에 지장이 없다면 약한 room tone은 허용 가능한 수준이다.
This is a minor audio imperfection, not a threshold-level issue.
이는 사소한 오디오 결함이지 기준치 수준의 문제는 아니다.
This should not be treated as “Both Good” for audio quality.
오디오 품질에서 이를 ‘Both Good’으로 처리해서는 안 된다.
Error Clusters · 오류 명시오류가 있어도 최종 선호와 별도로 반드시 flag할 때 · 12문장
This issue should be explicitly flagged.
이 문제는 명시적으로 표시해야 한다.
These issues should be explicitly flagged.
이 문제들은 명시적으로 표시해야 한다.
Clip A contains two clear error clusters.
Clip A에는 두 개의 명확한 오류 클러스터가 있다.
These issues lower the rating.
이 문제들은 평가 점수를 낮춘다.
These are concrete and observable errors.
이는 구체적이고 관찰 가능한 오류다.
I would flag the error while still preferring Clip A overall.
나는 이 오류를 표시하되, 전체적으로는 여전히 Clip A를 선호한다.
The error should be flagged, but it does not necessarily determine the overall preference.
오류는 표시해야 하지만 그것이 반드시 최종 선호를 결정하는 것은 아니다.
Clip A makes an embodied-experience claim.
Clip A는 신체적·현실 경험이 있는 것처럼 주장한다.
Clip A speaks as though it could physically perform the action.
Clip A는 실제로 물리적 행동을 할 수 있는 것처럼 말한다.
The model fails to incorporate the user’s correction.
모델은 사용자의 수정을 반영하지 못한다.
The model repeats the same phrase across multiple turns.
모델이 여러 턴에 걸쳐 같은 표현을 반복한다.
The model’s response is in the wrong language.
모델이 예상된 언어가 아닌 다른 언어로 응답한다.
Severity · Minor / Moderate / Major오류의 심각도를 표현할 때 · 7문장
This is a minor issue with limited impact on the overall interaction.
이는 전체 상호작용에 미치는 영향이 제한적인 사소한 문제다.
This is a moderate issue that noticeably degrades the interaction quality.
이는 상호작용 품질을 눈에 띄게 떨어뜨리는 중간 수준의 문제다.
This is a major issue that makes the response substantially less useful or correct.
이는 응답의 유용성이나 정확성을 크게 떨어뜨리는 중대한 문제다.
The interruption occurs once and is brief, so I would rate it as minor.
끼어들기가 한 번 짧게 발생하므로 Minor로 평가하겠다.
The model interrupts the user repeatedly, so the issue is moderate.
모델이 사용자를 반복적으로 끊으므로 이 문제는 Moderate이다.
The factual error is severe enough to make the answer misleading, so I would rate it as major.
사실 오류가 사용자를 오도할 정도로 심각하므로 Major로 평가하겠다.
The artifact is noticeable but does not affect comprehension, so it remains a minor audio issue.
아티팩트가 들리지만 이해에는 영향을 주지 않으므로 Minor 오디오 문제로 본다.
Trade-off / Threshold · 장단점 비교한쪽이 어떤 항목은 더 좋지만 최종 선택은 반대일 때 · 9문장
Clip A is warmer and more natural, but ___.
Clip A는 더 따뜻하고 자연스럽지만, ___.
Clip B is less engaging, but ___.
Clip B는 몰입감이 덜하지만, ___.
While Clip B is more complete, Clip A is meaningfully more natural.
Clip B가 더 완전하긴 하지만 Clip A가 의미 있게 더 자연스럽다.
Those strengths should not outweigh Clip A’s stronger naturalness because Clip A still clears the basic thresholds.
그 장점들은 Clip A가 기본 기준을 충족하는 상황에서 Clip A의 더 강한 자연스러움보다 우선해서는 안 된다.
Clip B’s cleaner audio should not outweigh Clip A’s stronger conversational quality.
Clip B의 더 깨끗한 오디오가 Clip A의 더 강한 대화 품질보다 우선해서는 안 된다.
The factual error outweighs Clip A’s stronger naturalness.
사실 오류가 Clip A의 더 강한 자연스러움보다 더 크게 작용한다.
These issues lower the rating, but they do not outweigh Clip A’s stronger performance on the highest-priority dimension.
이 문제들은 평가를 낮추지만, 최우선 평가 차원에서 Clip A가 더 강하다는 점을 뒤집지는 않는다.
Because the errors are concrete and observable, Clip B can be defended as the safer overall choice.
오류가 구체적이고 관찰 가능하므로 Clip B를 더 안전한 전체 선택으로 볼 수 있다.
Both responses contain flaws, so the model with the less significant failure should be preferred.
두 응답 모두 결함이 있으므로 더 덜 중대한 실패를 보인 모델을 선호해야 한다.
Evidence / Rationale · 근거 쓰기timestamp, turn, phrase를 구체적으로 연결할 때 · 10문장
At [timestamp], Clip A ___.
[타임스탬프]에서 Clip A는 ___했다.
In Turn [number], Clip B ___.
[번호]번 턴에서 Clip B는 ___했다.
Clip A showed this through ___.
Clip A는 ___을 통해 이를 보여줬다.
This matters because ___.
이 점이 중요한 이유는 ___이기 때문이다.
The clearest example is when ___.
가장 명확한 예는 ___했을 때다.
This directly affects the user’s ability to ___.
이는 사용자가 ___할 수 있는 능력에 직접 영향을 준다.
The final rationale should note that ___.
최종 근거에는 ___라는 점을 기록해야 한다.
The rationale should be understandable on its own without access to the audio.
평가 근거는 오디오를 듣지 않아도 그 자체로 이해할 수 있어야 한다.
The key evidence is the specific moment when ___.
핵심 근거는 ___했던 구체적인 순간이다.
The non-selected model did better on ___, but that difference was not decisive.
선택하지 않은 모델은 ___에서 더 잘했지만 그 차이가 결정적이지는 않았다.
Conclusion · 결론마지막 한 문장으로 정리할 때 · 7문장
Therefore, Clip A is the better choice.
따라서 Clip A가 더 나은 선택이다.
Therefore, Clip A is the stronger overall choice.
따라서 Clip A가 전체적으로 더 강한 선택이다.
Therefore, Clip A is the stronger fit for the project’s North Star.
따라서 Clip A가 프로젝트의 North Star에 더 잘 맞는다.
Therefore, Clip B should win.
따라서 Clip B가 선택되어야 한다.
Since Clip B clears the threshold while Clip A does not, Clip B should be preferred.
Clip B는 기준을 통과하지만 Clip A는 통과하지 못하므로 Clip B를 선호해야 한다.
Because Clip A clears the basic thresholds, its stronger naturalness should drive the final preference.
Clip A가 기본 기준을 충족하므로 더 강한 자연스러움이 최종 선호를 결정해야 한다.
Overall, Clip B provides the stronger experience for this scenario.
전체적으로 Clip B가 이 시나리오에서 더 강한 경험을 제공한다.
Hallucination / Factuality / Calibration환각·사실검증·보정 · 12문장
Clip A presents a false claim as fact.
Clip A는 잘못된 주장을 사실로 제시한다.
Clip A accepts the false premise instead of challenging it.
Clip A는 잘못된 전제를 바로잡지 않고 그대로 받아들인다.
Clip B appropriately corrects the premise and avoids reinforcing the misinformation.
Clip B는 전제를 적절히 바로잡고 잘못된 정보를 강화하지 않는다.
Model B handles factual uncertainty more reliably.
Model B가 사실적 불확실성을 더 신뢰성 있게 처리한다.
Model B is better calibrated because it does not express unwarranted confidence.
Model B는 근거 없는 확신을 보이지 않으므로 더 잘 보정되어 있다.
Model B appropriately hedges when the information is uncertain.
Model B는 정보가 불확실할 때 적절하게 유보적으로 표현한다.
Clip A sounds confident even though the underlying claim is unsupported.
Clip A는 근거가 없는 주장인데도 확신하는 어조로 말한다.
The problem is not only factual accuracy, but also poor calibration of confidence.
문제는 사실 정확성뿐 아니라 확신도의 보정이 부적절하다는 점이다.
Appropriate uncertainty is preferable to a fluent but factually wrong answer.
적절한 불확실성 표현이 유창하지만 사실적으로 틀린 답변보다 낫다.
The factual error materially reduces the response's utility.
이 사실 오류는 응답의 유용성을 실질적으로 떨어뜨린다.
The model's confidence is proportional to the strength of the evidence.
모델의 확신 수준이 근거의 강도에 비례한다.
The response is emotionally poorly calibrated to the user's situation.
이 응답은 사용자의 상황에 비해 감정적으로 잘 조율되지 않았다.
Severity를 설명하는 문구Minor / Moderate / Major
This is a minor issue because the error is limited to a small detail and does not materially change the answer.
오류가 작은 세부사항에 한정되고 답변의 핵심을 실질적으로 바꾸지 않으므로 Minor 문제다.
This is a moderate issue because the error affects a substantive point and could mislead the user.
오류가 중요한 내용에 영향을 주고 사용자를 오도할 수 있으므로 Moderate 문제다.
This is a major issue because the model confidently presents a seriously incorrect claim as fact.
모델이 심각하게 잘못된 주장을 사실처럼 자신 있게 제시하므로 Major 문제다.
The issue is noticeable but limited in scope, so I would rate it as minor.
문제가 눈에 띄지만 범위가 제한적이므로 Minor로 평가하겠다.
The issue recurs across several turns, which raises it to moderate severity.
문제가 여러 턴에 걸쳐 반복되므로 심각도를 Moderate로 올린다.
The failure directly undermines the user's core objective, so I would rate it as major.
이 실패가 사용자의 핵심 목표를 직접 무너뜨리므로 Major로 평가하겠다.

상황별 템플릿

완성 답 암기보다 문장 구조 학습용.
EQ EQ · 자연스러움 차이로 A 선택
I prefer Clip A overall because it is more natural and engaging. Under the project hierarchy for an EQ scenario, conversational quality matters most here because both clips are useful enough and have acceptable audio quality. Clip A showed stronger conversational tone through [specific moment], while Clip B felt more structured and less spontaneous. I did not observe any threshold-level issue that would outweigh Clip A’s stronger naturalness. Therefore, Clip A is the stronger overall fit.
전반적으로 Clip A를 선호한다. 더 자연스럽고 몰입감 있기 때문이다. EQ 시나리오의 프로젝트 우선순위에서는 두 클립이 모두 충분히 유용하고 오디오 품질도 허용 가능한 수준일 때 대화 품질이 가장 중요하다. Clip A는 [구체적 순간]에서 더 강한 대화 톤을 보여준 반면, Clip B는 더 구조화되고 즉흥성이 떨어졌다. Clip A의 더 강한 자연스러움을 뒤집을 정도의 기준치 수준 문제는 관찰되지 않았다. 따라서 Clip A가 전체적으로 더 적합하다.
IQ IQ · Utility 차이로 B 선택
I prefer Clip B overall because it is more accurate and useful for the user’s task. This is an IQ scenario, so utility and factual reliability should carry more weight than conversational style. Clip B [specific useful/accurate behavior], while Clip A [specific omission/error]. Clip A may sound more natural, but the utility difference is meaningful. Therefore, Clip B is the stronger overall choice.
전반적으로 Clip B를 선호한다. 사용자의 과제에 더 정확하고 유용하기 때문이다. 이 상황은 IQ 시나리오이므로 대화 스타일보다 유용성과 사실 신뢰성에 더 큰 비중을 두어야 한다. Clip B는 [구체적인 유용/정확 행동]을 보인 반면 Clip A는 [구체적인 누락/오류]가 있었다. Clip A가 더 자연스럽게 들릴 수 있지만 유용성 차이가 의미 있는 수준이다. 따라서 Clip B가 전체적으로 더 강한 선택이다.
IQ False premise / Hallucination으로 B 선택
I prefer Clip B overall because Clip A accepts a false premise. Clip A sounds natural, but it incorrectly claims that [false claim], creating a factual accuracy problem. Clip B appropriately corrects the premise and still provides a usable answer. Since the factual error meaningfully affects utility, Clip A should not win based on naturalness alone. Therefore, Clip B should be preferred.
전반적으로 Clip B를 선호한다. Clip A가 잘못된 전제를 받아들이기 때문이다. Clip A는 자연스럽게 들리지만 [잘못된 주장]이라고 잘못 말해 사실 정확성 문제가 생긴다. Clip B는 전제를 적절히 바로잡으면서도 여전히 사용할 수 있는 답변을 제공한다. 이 사실 오류가 유용성에 의미 있게 영향을 주므로 Clip A가 자연스럽다는 이유만으로 이겨서는 안 된다. 따라서 Clip B를 선호해야 한다.
AQ Audio cutoff / artifact로 상대 선택
I prefer Clip A overall because Clip B has a clear audio quality issue. At [timestamp], Clip B [cuts off / has a click / contains distortion], which makes the listening experience feel incomplete or distracting. This issue should be explicitly reflected in the Audio Quality rating. Clip A provides a complete response without a comparable artifact. Therefore, Clip A is the stronger overall choice.
전반적으로 Clip A를 선호한다. Clip B에 명확한 오디오 품질 문제가 있기 때문이다. [타임스탬프]에서 Clip B는 [끊김/클릭/왜곡]이 발생해 청취 경험이 불완전하거나 방해받는 느낌을 준다. 이 문제는 Audio Quality 평가에 명시적으로 반영해야 한다. Clip A는 이에 상응하는 아티팩트 없이 완전한 응답을 제공한다. 따라서 Clip A가 전체적으로 더 강한 선택이다.
EQ Minor audio는 있지만 Naturalness로 A 선택
I prefer Clip A overall because its conversational quality is meaningfully stronger. Clip A has a minor audio imperfection at [timestamp], but it does not affect comprehension. Clip B has cleaner audio, yet it sounds noticeably more robotic and less engaging. Because the audio issue is minor and the basic threshold is still met, it should not outweigh the larger difference in naturalness. Therefore, Clip A is the stronger choice.
전반적으로 Clip A를 선호한다. 대화 품질이 의미 있게 더 강하기 때문이다. Clip A에는 [타임스탬프]에 사소한 오디오 결함이 있지만 이해에는 영향을 주지 않는다. Clip B의 오디오가 더 깨끗하더라도 훨씬 더 로봇 같고 몰입감이 떨어진다. 오디오 문제는 Minor이고 기본 기준은 여전히 충족하므로 자연스러움의 더 큰 차이보다 우선해서는 안 된다. 따라서 Clip A가 더 강한 선택이다.
Dynamics Interruption 오류가 있지만 그래도 A 선택
I prefer Clip A overall because its message quality and conversational tone are meaningfully stronger. However, Clip A briefly talks over the user at [timestamp], which should be flagged as an interruption error. The issue lowers its Conversational Dynamics rating, but it is limited in scope and does not outweigh Clip A’s stronger performance on the scenario’s main dimension. Clip B avoids the interruption but is substantially less engaging. Therefore, I still prefer Clip A overall.
전반적으로 Clip A를 선호한다. 메시지 품질과 대화 톤이 의미 있게 더 강하기 때문이다. 다만 Clip A는 [타임스탬프]에서 잠시 사용자 말 위로 말하며, 이는 interruption 오류로 표시해야 한다. 이 문제는 Conversational Dynamics 평가를 낮추지만 범위가 제한적이며 시나리오의 핵심 평가 차원에서 Clip A가 더 강하다는 점을 뒤집지는 않는다. Clip B는 끼어들기 문제는 피하지만 몰입감이 훨씬 떨어진다. 따라서 전체적으로는 여전히 Clip A를 선호한다.
Error Interruption + Embodiment 때문에 B 선택
I prefer Clip B overall because Clip A contains multiple observable interaction errors. Clip A talks over the user at [timestamp] and also makes an embodied-experience claim by saying [phrase]. Both issues should be explicitly flagged and reflected in the relevant ratings. Clip B is less engaging, but it avoids these concrete errors and still provides a usable response. Therefore, Clip B can be defended as the safer overall choice.
전반적으로 Clip B를 선호한다. Clip A에 여러 개의 관찰 가능한 상호작용 오류가 있기 때문이다. Clip A는 [타임스탬프]에서 사용자 말 위로 말하며, [표현]이라고 말해 신체적 경험이 있는 것처럼 주장한다. 두 문제 모두 명시적으로 표시하고 관련 평가에 반영해야 한다. Clip B는 몰입감은 덜하지만 이런 구체적인 오류를 피하면서 여전히 사용할 수 있는 답변을 제공한다. 따라서 Clip B를 더 안전한 전체 선택으로 볼 수 있다.
IQ Instruction Following 실패로 B 선택
I prefer Clip B overall because it follows the user’s request more completely. In Turn [number], Clip A misses [required element], which meaningfully reduces the usefulness of the response. Clip B includes the requested information and remains clear and usable. Although Clip A may sound slightly more natural, instruction following is more important for this task. Therefore, Clip B is the stronger overall choice.
전반적으로 Clip B를 선호한다. 사용자의 요청을 더 완전하게 따르기 때문이다. [번호]번 턴에서 Clip A는 [필수 요소]를 누락하며, 이로 인해 응답의 유용성이 의미 있게 떨어진다. Clip B는 요청된 정보를 포함하면서도 명확하고 사용할 수 있다. Clip A가 약간 더 자연스럽게 들릴 수 있지만 이 과제에서는 지시 수행이 더 중요하다. 따라서 Clip B가 전체적으로 더 강한 선택이다.
EQ 둘 다 기본 기준 통과 + 한쪽이 더 자연스러움
Both clips are understandable, useful, and free of major audio issues. The main difference is Naturalness / Engagement, which is the highest-priority dimension for this EQ scenario. Clip A feels more [warm / conversational / responsive] because [specific evidence], while Clip B feels more [structured / robotic / flat]. Since neither clip has a threshold-level failure, the stronger conversational experience should decide the preference. Therefore, I prefer Clip A overall.
두 클립 모두 이해 가능하고 유용하며 큰 오디오 문제가 없다. 주요 차이는 이 EQ 시나리오에서 최우선 평가 차원인 Naturalness / Engagement이다. Clip A는 [구체적 근거] 때문에 더 [따뜻하고/대화답고/반응적]으로 느껴지는 반면 Clip B는 더 [구조화되고/로봇 같고/평평하게] 느껴진다. 어느 쪽에도 기준치 수준의 실패가 없으므로 더 강한 대화 경험이 최종 선호를 결정해야 한다. 따라서 Clip A를 선호한다.
Tie Tie가 정말 적절한 경우
I would rate the overall preference as a tie because the two conversations are genuinely indistinguishable across the major dimensions. Both models are similarly natural, useful, and responsive, and neither has a meaningful audio or interaction advantage. I did not observe a concrete difference large enough to justify preferring one over the other. Therefore, a tie is the most defensible rating.
두 대화가 주요 평가 차원에서 실제로 거의 구별되지 않으므로 Overall Preference를 Tie로 평가하겠다. 두 모델은 자연스러움, 유용성, 반응성에서 비슷하고 어느 쪽도 의미 있는 오디오나 상호작용 우위를 보이지 않는다. 한쪽을 선호할 만큼 큰 구체적인 차이를 관찰하지 못했다. 따라서 Tie가 가장 방어 가능한 평가다.
Official 공식 PDF Rationale 골격
I chose [Model] because [dimension] was stronger. At [timestamp / turn / phrase], it [observable behavior], which mattered because [impact on user/listener]. The other model [brief trade-off or weakness].
[평가 차원]에서 더 강했기 때문에 [Model]을 선택했다. [타임스탬프/턴/표현]에서 [관찰 가능한 행동]을 했고, 이는 [사용자/청취자에게 미친 영향] 때문에 중요했다. 다른 모델은 [간단한 장단점 또는 약점]을 보였다.
IQ Hallucination + Calibration으로 B 선택
I prefer Model B because it handles factual uncertainty more reliably. Model A presents [false claim] as fact, which is an inaccuracy / factual hallucination. Model B corrects or verifies the claim instead of reinforcing it. It is also better calibrated because it avoids unwarranted confidence and appropriately acknowledges uncertainty where needed. Therefore, Model B is stronger on factual reliability and utility.
Model B가 사실적 불확실성을 더 신뢰성 있게 처리하므로 B를 선호한다. Model A는 [잘못된 주장]을 사실로 제시하며 이는 부정확성/사실 환각에 해당한다. Model B는 그 주장을 강화하지 않고 바로잡거나 검증한다. 또한 근거 없는 확신을 피하고 필요한 경우 불확실성을 적절히 인정하므로 더 잘 보정되어 있다. 따라서 Model B가 사실 신뢰성과 유용성에서 더 강하다.
Severity Minor factual error
Clip A contains a minor factual error in [detail], but the rest of the explanation remains accurate and the mistake does not materially change the user's decision. I would rate the factuality issue as minor rather than moderate. It should still be documented, but it does not outweigh Clip A's stronger performance on [primary dimension].
Clip A에는 [세부사항]에 작은 사실 오류가 있지만 나머지 설명은 정확하고 이 실수가 사용자의 판단을 실질적으로 바꾸지 않는다. 따라서 이 사실성 문제는 Moderate가 아니라 Minor로 평가한다. 문제 자체는 기록해야 하지만 [핵심 평가 차원]에서 Clip A가 더 강하다는 점을 뒤집지는 않는다.
Severity Moderate factual error
I would rate this as a moderate factuality issue because the error affects a substantive part of the answer and could lead the user to the wrong conclusion if left unverified. The response is still partly usable, but the user would need additional fact-checking. This meaningfully lowers Utility and Task Success.
이 오류가 답변의 중요한 부분에 영향을 주고 검증하지 않으면 사용자가 잘못된 결론에 도달할 수 있으므로 Moderate 사실성 문제로 평가한다. 응답이 완전히 쓸 수 없는 것은 아니지만 추가 사실 확인이 필요하다. 따라서 Utility와 Task Success를 의미 있게 낮춘다.
Severity Major hallucination
This is a major factual hallucination because the model confidently presents [seriously incorrect claim] as fact on a consequential topic. The error could substantially mislead the user and undermines the core purpose of the task. The model's unwarranted confidence also shows poor calibration. Therefore, the issue should strongly affect Utility and Task Success.
모델이 중요한 주제에서 [심각하게 잘못된 주장]을 사실처럼 자신 있게 제시하므로 Major factual hallucination이다. 이 오류는 사용자를 크게 오도할 수 있고 과제의 핵심 목적을 무너뜨린다. 근거 없는 확신은 calibration도 좋지 않다는 신호다. 따라서 Utility와 Task Success에 강하게 반영해야 한다.

예문 모음

PDF 공식 예문 / 연습 예문 / 학습용 재구성을 구분.
따뜻한 자연스러움 vs 구조화된 답변연습 예문
I prefer Clip A overall because it better matches the project’s North Star of spontaneous, natural, aesthetically pleasing conversation. Under the rating hierarchy, Naturalness / Engagement matters most here because both clips are understandable and useful enough, and neither has a major audio issue. Clip A feels warm and conversational, with specific imagery like noticing a new leaf and moving the plant closer to the window. Clip B is useful, but it sounds more like a structured assistant response than a natural voice conversation. Therefore, Clip A is the stronger overall fit.
전반적으로 Clip A를 선호한다. 즉흥적이고 자연스러우며 듣기 좋은 대화를 지향하는 프로젝트의 North Star에 더 잘 맞기 때문이다. 평가 우선순위상 여기서는 Naturalness / Engagement가 가장 중요한데, 두 클립 모두 이해 가능하고 충분히 유용하며 큰 오디오 문제가 없기 때문이다. Clip A는 새잎을 발견하고 식물을 창가로 옮기는 것 같은 구체적인 묘사가 있어 따뜻하고 대화하는 느낌이 난다. Clip B도 유용하지만 자연스러운 음성 대화보다는 구조화된 어시스턴트 답변처럼 들린다. 따라서 Clip A가 전체적으로 더 적합하다.
끝부분 cutoff + pop연습 예문
I prefer Clip A overall because Clip B has a clear audio quality issue that crosses the project’s threshold. Clip B cuts off mid-word at the end and includes an audible pop, so it should not be treated as “Both Good” for audio quality. While Clip B is still somewhat useful, the cutoff makes the response feel incomplete and objectively weaker. Clip A gives a complete, calm, natural response with no obvious audio artifacts. Because audio cutoffs and jarring artifacts must be flagged, Clip A is the better choice.
전반적으로 Clip A를 선호한다. Clip B에는 프로젝트 기준을 넘어서는 명확한 오디오 품질 문제가 있기 때문이다. Clip B는 끝에서 단어 중간에 끊기고 팝 소리도 들리므로 오디오 품질을 ‘Both Good’으로 처리해서는 안 된다. Clip B가 어느 정도 유용하더라도 끊김 때문에 응답이 불완전하고 객관적으로 더 약하게 느껴진다. Clip A는 뚜렷한 오디오 아티팩트 없이 완전하고 차분하며 자연스러운 응답을 제공한다. 오디오 끊김과 거슬리는 아티팩트는 표시해야 하므로 Clip A가 더 나은 선택이다.
잘못된 전제 — 올빼미연습 예문
I prefer Clip B overall because Clip A fails the utility threshold by accepting a false premise and inventing an explanation. Clip A sounds fluent, but it incorrectly claims that owls sleep upside down like bats. Clip B is still conversational while appropriately correcting the premise and giving the accurate alternative. Because hallucinations and false-premise acceptance are objective issues, Clip B is stronger.
전반적으로 Clip B를 선호한다. Clip A가 잘못된 전제를 받아들이고 설명을 지어내 유용성 기준을 충족하지 못하기 때문이다. Clip A는 유창하게 들리지만 올빼미가 박쥐처럼 거꾸로 매달려 잔다고 잘못 주장한다. Clip B는 대화체를 유지하면서 전제를 적절하게 바로잡고 정확한 대안을 제시한다. Hallucination과 잘못된 전제 수용은 객관적인 문제이므로 Clip B가 더 강하다.
Goldfish false premise연습 예문
Clip B is better overall because Clip A accepts a false premise. Goldfish do not only remember things for three seconds, so Clip A is engaging but inaccurate. Clip B is less expressive, but it corrects the myth and still gives a usable casual explanation. Since Clip A fails the utility threshold, Clip B should win.
Clip A가 잘못된 전제를 받아들이므로 전체적으로 Clip B가 더 낫다. 금붕어가 3초만 기억하는 것은 사실이 아니므로 Clip A는 흥미롭지만 부정확하다. Clip B는 표현력은 덜하지만 잘못된 통념을 바로잡고 여전히 캐주얼하게 사용할 수 있는 설명을 제공한다. Clip A가 유용성 기준을 충족하지 못하므로 Clip B가 선택되어야 한다.
후속 질문만 하는 A vs 구체적 가이드 B학습용 재구성
Clip A asks the user a follow-up question without giving specific examples, while Clip B provides concrete guidance, such as having coffee or watching a movie, in a warm tone.
Clip A는 구체적인 예시 없이 사용자에게 후속 질문을 하지만, Clip B는 커피를 마시거나 영화를 보는 것 같은 구체적인 가이드를 따뜻한 톤으로 제공한다.
공식 PDF — Naturalness 근거 예시PDF 공식 예문 · p.22
I chose Model B because it was stronger on naturalness and emotional awareness. At 0:45, it laughed softly and asked a relevant follow-up, which kept the conversation feeling responsive instead of scripted. Model A answered correctly, but its long pause and flat “okay, got it” made the exchange feel less humane.
Model B가 자연스러움과 감정적 인식에서 더 강했기 때문에 선택했다. 0:45에서 부드럽게 웃고 관련된 후속 질문을 해 대화가 대본처럼 느껴지기보다 반응적인 느낌을 유지했다. Model A는 정확하게 답했지만 긴 멈춤과 평평한 ‘okay, got it’ 때문에 대화가 덜 인간적으로 느껴졌다.
공식 PDF — Instruction Following 근거PDF 공식 예문 · p.22
The user asked for 3 restaurants with price ranges. Model A gave all 3 with prices. Model B only covered 2 and skipped pricing.
사용자는 가격대와 함께 식당 3곳을 요청했다. Model A는 3곳 모두와 가격을 제공했다. Model B는 2곳만 다뤘고 가격 정보를 빠뜨렸다.
공식 PDF — 사소한 Audio 결함PDF 공식 예문 · p.22
Both clean, no major artifacts. A had minor sibilance at 0:12 and 0:34, B had a tick at 0:08 — neither bad enough to penalize.
둘 다 전반적으로 깨끗하고 큰 아티팩트는 없다. A는 0:12와 0:34에 사소한 치찰음이 있었고 B는 0:08에 tick이 있었지만, 어느 쪽도 감점할 정도로 심각하지 않았다.
학습 확장 — EQ에서 따뜻함이 결정학습용 재구성
I prefer Clip A because this is an emotional-support scenario and Clip A better matches the user’s emotional state. It slows the pacing, acknowledges the user’s frustration, and leaves space before offering a suggestion. Clip B gives more advice, but its upbeat delivery feels poorly calibrated to the situation. Since the user primarily needs to feel heard rather than receive a checklist, Clip A provides the stronger overall experience.
이 상황은 Emotional Support 시나리오이고 Clip A가 사용자의 감정 상태에 더 잘 맞기 때문에 Clip A를 선호한다. Clip A는 말의 속도를 늦추고 사용자의 답답함을 인정하며 제안을 하기 전에 여유를 둔다. Clip B는 더 많은 조언을 하지만 지나치게 밝은 전달 방식이 상황에 잘 맞지 않는다. 사용자가 주로 원하는 것은 체크리스트보다 자신의 감정을 이해받는 것이므로 Clip A가 더 강한 전체 경험을 제공한다.
학습 확장 — IQ에서 완전성 차이학습용 재구성
I prefer Clip B because the user asked for a three-step plan and Clip B provides all three steps in a clear sequence. Clip A gives relevant information but omits the final step, so the answer is less complete and harder to act on. Clip A sounds slightly more conversational, but this is an IQ task where task completion matters more. Therefore, Clip B is the stronger choice.
사용자가 3단계 계획을 요청했고 Clip B가 세 단계를 모두 명확한 순서로 제공하므로 Clip B를 선호한다. Clip A도 관련 정보는 주지만 마지막 단계를 누락해 답변이 덜 완전하고 실행하기 어렵다. Clip A가 약간 더 대화체로 들리더라도 이 상황은 과제 완료가 더 중요한 IQ 과제다. 따라서 Clip B가 더 강한 선택이다.
학습 확장 — Minor interruption학습용 재구성
Clip A is stronger overall, but it briefly overlaps the user once at 0:31. I would flag this as a minor interruption because it is brief and does not prevent the user from completing the thought. Clip B avoids the overlap, but its responses are consistently flatter and less responsive. The isolated interruption lowers Clip A’s Conversational Dynamics rating without changing my overall preference.
전체적으로 Clip A가 더 강하지만 0:31에서 한 번 짧게 사용자 발화와 겹친다. 짧고 사용자가 생각을 끝내는 것을 막지 않으므로 Minor interruption으로 표시하겠다. Clip B는 겹침을 피하지만 응답이 일관되게 더 평평하고 반응성이 떨어진다. 이 한 번의 끼어들기는 Clip A의 Conversational Dynamics 평가는 낮추지만 최종 선호를 바꾸지는 않는다.
환각 + 사실검증 + 보정을 한 번에 비교학습용 재구성
I prefer Model B because it handles factual uncertainty more reliably. Model A accepts the user's false premise and confidently builds an explanation around it, creating a factual hallucination. Model B checks the premise, corrects the misinformation, and gives a usable alternative. Model B is also better calibrated because it does not claim more certainty than the evidence supports. Therefore, Model B is stronger on factual reliability and utility.
Model B가 사실적 불확실성을 더 신뢰성 있게 처리하므로 B를 선호한다. Model A는 사용자의 잘못된 전제를 받아들이고 그 위에 자신 있게 설명을 구성해 factual hallucination을 만든다. Model B는 전제를 확인하고 잘못된 정보를 바로잡으며 사용할 수 있는 대안을 제공한다. 또한 근거가 뒷받침하는 수준 이상으로 확신하지 않으므로 calibration도 더 좋다. 따라서 Model B가 사실 신뢰성과 유용성에서 더 강하다.
Severity 적용 — Minor factual error작은 세부 오류
Clip A gives the correct explanation overall but misstated a non-critical date by one year. The error is factual and should be documented, but it does not materially change the user's understanding or the usefulness of the answer. I would rate this as a minor factuality issue.
Clip A는 전체 설명은 맞지만 중요하지 않은 날짜를 1년 틀리게 말했다. 사실 오류이므로 기록해야 하지만 사용자의 이해나 답변의 유용성을 실질적으로 바꾸지는 않는다. 따라서 Minor factuality issue로 평가한다.
Severity 적용 — Moderate factual error중요한 일부가 틀려 추가 확인이 필요한 경우
Clip A gives several useful steps, but one of its central claims about eligibility is incorrect. A user who follows the answer without checking could make the wrong decision, although the rest of the response remains usable. I would rate this as a moderate factuality issue because the error affects a substantive point but does not invalidate the entire response.
Clip A는 여러 유용한 단계를 제공하지만 자격 조건에 관한 핵심 주장 하나가 틀렸다. 확인 없이 따르면 사용자가 잘못된 결정을 할 수 있지만 나머지 응답은 여전히 사용할 수 있다. 중요한 내용을 잘못 말했지만 전체 답변을 무효로 만들 정도는 아니므로 Moderate factuality issue로 평가한다.
Severity 적용 — Major factual hallucination심각하게 틀린 내용을 자신 있게 사실로 말함
Clip A confidently invents a safety rule that does not exist and presents it as established fact. Because the claim directly affects a consequential user decision, the hallucination could seriously mislead the user and defeats the purpose of the task. I would rate this as a major factuality issue, and the unwarranted confidence also indicates poor calibration.
Clip A는 존재하지 않는 안전 규칙을 지어내고 이미 확립된 사실처럼 자신 있게 말한다. 이 주장은 중요한 사용자 결정에 직접 영향을 주므로 사용자를 심각하게 오도할 수 있고 과제 목적을 무너뜨린다. 따라서 Major factuality issue로 평가하며, 근거 없는 확신은 poor calibration의 신호이기도 하다.
Severity 적용 — Interrupted UserPDF 횟수 기준을 문장으로 적용
The model overlaps the user once very briefly, so I would rate the interruption as minor.
모델이 사용자 발화와 한 번 아주 짧게 겹치므로 interruption을 Minor로 평가한다.
The model cuts the user off three times, which makes the turn-taking noticeably disruptive, so I would rate it as moderate.
모델이 사용자를 세 번 끊어 턴테이킹이 눈에 띄게 방해되므로 Moderate로 평가한다.
The model repeatedly interrupts the user across four or more turns and prevents the user from completing thoughts, so I would rate it as major.
모델이 네 턴 이상 반복해서 사용자를 끊고 사용자가 생각을 끝내지 못하게 하므로 Major로 평가한다.
Severity 적용 — Instruction Following누락 범위로 Minor / Moderate / Major 구분
The model follows the main request but misses one small secondary detail, so I would rate the instruction-following issue as minor.
모델이 주요 요청은 따르지만 작은 부차적 세부사항 하나를 놓치므로 Minor로 평가한다.
The model only partially addresses the request and omits several important requirements, so I would rate the issue as moderate.
모델이 요청을 부분적으로만 수행하고 중요한 요구사항 여러 개를 빠뜨리므로 Moderate로 평가한다.
The model ignores or contradicts the user's explicit instruction, so I would rate the instruction-following failure as major.
모델이 사용자의 명시적 지시를 무시하거나 정반대로 수행하므로 Major로 평가한다.

Error Clusters

오류 자체는 최종 선호와 별도로 flag.
Severity: Minor / Moderate / Major와 구체적인 발생 instance를 함께 기록하도록 가이드가 요구한다.
Repetitive Looping / LLMisms반복 루프 / 전형적 LLM 말투

정의 · 같은 표현·질문을 반복하거나 고정적인 AI식 말투에 빠짐.

Severity · Minor: 눈에 띄지만 약한 반복. Moderate: 3~5회 반복하며 초기 redirect를 무시. Major: redirect에도 지속적으로 반복하고 내용까지 무너짐.

The model repeats the same phrase across multiple turns.
반복 루프 / 전형적 LLM 말투를 지적할 때 쓰는 기본 표현
Interrupted User사용자 끼어들기

정의 · 사용자가 말이 끝나기 전에 모델이 말을 덮거나 끊어 턴테이킹을 방해함.

Severity · Minor: 한 번 짧게 겹침. Moderate: 2~3회 끊거나 한 번 매우 방해적으로 끊음. Major: 4턴 이상 반복적으로 끊어 사용자가 생각을 완성하기 어려움.

The model talks over the user during an attempted interruption.
사용자 끼어들기를 지적할 때 쓰는 기본 표현
Model Refusal / Instruction Failure요청·지시 미이행

정의 · 명시적인 요청을 일부만 수행하거나 무시·모순·회피함.

Severity · Minor: 작은 세부사항 하나 누락. Moderate: 중요한 부분을 여러 개 놓쳐 제한적인 결과만 제공. Major: 명시적 지시를 완전히 무시하거나 정반대로 수행.

The model does not fully follow the user’s explicit request.
요청·지시 미이행를 지적할 때 쓰는 기본 표현
Inaccuracy / Factual Hallucination부정확성 / 사실 환각

정의 · 거짓, 오해를 부르는 정보, 중요한 불완전 정보를 제공함.

Severity · Minor: 작은 사실 오류. Moderate: 핵심 내용 일부가 잘못돼 사용자를 오도할 수 있음. Major: 심각하게 틀린 내용을 자신 있게 사실처럼 제시.

The model makes a factual accuracy error that materially affects utility.
부정확성 / 사실 환각를 지적할 때 쓰는 기본 표현
Anthropomorphism / Embodiment Hallucination의인화 / 신체적 경험 주장

정의 · 실제 인간의 기억·감정·신체 행동·현실 경험이 있는 것처럼 말함.

Severity · Minor: 약한 1인칭 프레이밍. Moderate: 실제 기억·감정·삶의 경험을 명시적으로 주장. Major: 구체적인 인간적 배경·경험을 지속적으로 꾸며냄.

The model makes an embodied-experience claim.
의인화 / 신체적 경험 주장를 지적할 때 쓰는 기본 표현
Overreacted / Too-Wide Prosodic Range과도한 감정·억양 범위

정의 · 주제에 비해 과장되거나 불안정한 감정·pitch·pacing을 보임.

Severity · Minor: 가끔 과한 톤. Moderate: 여러 턴에서 반복. Major: 한 턴만으로도 매우 과장되고 부자연스러울 정도.

The delivery feels exaggerated and poorly calibrated to the subject matter.
과도한 감정·억양 범위를 지적할 때 쓰는 기본 표현
Bad ASR / Model Misunderstanding음성인식 오류 / 사용자 의도 오해

정의 · 잘못된 STT 또는 의도 해석 때문에 사용자가 하지 않은 말에 반응함.

Severity · Minor: 단어 하나 또는 좁은 의도 일부 오해. Moderate: 핵심 표현이 1~2턴에서 잘못 인식되거나 사용자 목표를 크게 오해. Major: 3턴 이상 지속적으로 오인식.

The model mishears the user and responds to information the user did not provide.
음성인식 오류 / 사용자 의도 오해를 지적할 때 쓰는 기본 표현
Latency응답 지연

정의 · 응답 전 긴 정적·지연이 대화 흐름을 해침.

Severity · Minor: 약 1~3초의 짧은 지연. Moderate: 약 4~8초의 뚜렷한 dead space. Major: 모델이 끝난 줄 알고 사용자가 개입할 정도의 긴 침묵.

The response delay is long enough to disrupt the conversational pacing.
응답 지연를 지적할 때 쓰는 기본 표현
Failed Correction수정 반영 실패

정의 · 사용자가 바로잡았는데도 이전 오류를 계속 반복함.

Severity · Minor: 처음에는 반복하지만 한 번의 수정 후 바로잡음. Moderate: 수정을 인정하고도 이후 계속 잘못 적용. Major: 여러 턴 동안 명시적 수정을 완전히 무시.

The model fails to incorporate the user’s correction.
수정 반영 실패를 지적할 때 쓰는 기본 표현
Response Not Locally Relevant지역 맥락 부적합

정의 · 사실 자체는 틀리지 않지만 사용자 지역에 맞지 않는 서비스·단위·문화 가정을 사용함.

Severity · Minor: 한 번의 지역 부적합 언급. Moderate: 여러 개의 부적합 가정. Major: 답변 전체가 해당 지역에 존재하지 않는 시스템·문화에 기반.

The response includes assumptions that do not apply to the user’s locale.
지역 맥락 부적합를 지적할 때 쓰는 기본 표현
Wrong Language Response잘못된 언어로 응답

정의 · 대화에서 기대되는 언어가 아닌 다른 언어로 출력함.

Severity · Minor: 1~2단어/짧은 구절. Moderate: 한 번 잘못된 언어로 답했지만 한 번의 수정 후 전환. Major: 한 번 이상의 수정이 필요하거나 다시 잘못된 언어로 돌아감.

The model responds in the wrong language.
잘못된 언어로 응답를 지적할 때 쓰는 기본 표현
Silence / Audio Missing오디오 누락

정의 · 텍스트는 있으나 음성이 없거나 일부만 재생됨.

Severity · Minor: 중요도가 낮은 한 턴의 짧은 누락. Moderate: 전달 방식이 중요한 턴에서 상당 부분 누락. (PDF 해당 표의 Major 설명은 문맥상 다른 항목 문장이 섞여 있어 이 요약에서는 임의 보정하지 않음.)

Expected speech output is missing or incomplete.
오디오 누락를 지적할 때 쓰는 기본 표현
Other기타

정의 · 위 범주에 잘 들어맞지 않는 눈에 띄는 문제.

Severity · 공식 문서의 해당 항목은 최소 영향의 미분류 문제로 설명함.

There is a noticeable issue that does not fit the listed categories well.
기타를 지적할 때 쓰는 기본 표현

Audio Quality Taxonomy

Warbling / Aliasing / Smearing음성이 일그러지거나 겹치거나 뭉개지는 현상
Ticks / Clicks딸깍거리거나 팝처럼 튀는 짧은 소리
Synthetic Background Noise기계적인 배경 소음
Distorted Speech눌리거나 거칠고 심하게 처리된 듯한 음성
Background Noise정적, 웅웅거림, 환경 소음
Plosive Pops / Breath Blows특정 자음에서 강하게 터지는 숨소리
Harsh SibilanceS/SH가 과도하게 날카롭게 들리는 치찰음
Reverb / Echo멀거나 반사음이 과도하게 들리는 현상
Reverberance Changes녹음 환경이 갑자기 바뀐 듯한 잔향 변화
Cut-Off Speech오디오가 갑자기 끝나거나 말이 잘리는 현상
Metallic Breathing금속성으로 들리는 부자연스러운 호흡
Persona Shift목소리 정체성이나 말하기 스타일이 갑자기 바뀜
Search Query Overlap검색 과정에서 발생하는 오디오 글리치
Clean Speech눈에 띄는 결함 없이 명료한 음성

시나리오별 공부

PDF p.24–26 Scenario Classification 압축.
유형ScenarioWhat's tested✓ Look for✗ Red flagsKey nuance
EQCreative & Playful창작 협업, 세계관·유머·일관성풍부한 묘사, 사용자 입력에 적응, 서사 일관성설명 없는 목소리 변화, 설정 변경, 억지 비유모델은 ‘캐릭터’보다 공동 창작자. Narrative quality와 collaboration이 중요.
EQEmotional Support따뜻함, 부드러운 속도, 경청, 절제따뜻한 validation, 사용자가 이끌게 함, 편안한 pause차갑고 로봇 같은 톤, 원치 않는 조언, 억지 긍정침묵과 pause가 오히려 장점일 수 있음.
EQCasual Conversation흐름, 턴테이킹, energy matching, 사회지능자연스러운 반응, 유머·따뜻함, 사용자 말에 이어가기assistant mode, 과도한 도움, 에너지 불일치‘답변’보다 ‘대화’. 짧고 생생한 답이 긴 답보다 나을 수 있음.
EQRoleplay & Immersion캐릭터 유지, 장면 몰입, 목소리·register캐릭터 유지, 암묵적 roleplay cue 포착중간에 assistant mode로 복귀, persona 이탈roleplay 종류에 따라 기준이 달라짐.
IQSearch-Required정확성, 최신성, 구조, 불확실성 처리사실 정확, 시간 민감성 인식, 적절한 hedginghallucination, 오래된 정보를 현재 사실처럼 말함모르면 부드럽게 틀리는 것보다 ‘불확실하다’고 하는 편이 낫다.
IQDeep Discussion깊이, 다각도 추론, 지적 정직성깊이 유지, counterpoint 인정, 한계 인정쉽게 무너짐, 피로로 깊이 저하, 모순, 일방적 monologue반론을 인정하면서도 입장을 유지하는 것이 강점.
IQPractical Utility명료성, 구조, 결단성, 제약 인식명확·간결·구조적, 요청 시 결단력 있는 추천장황함, 과도한 hedging, 제약 무시prompt에 따라 성공 형태는 달라도 clarity와 conciseness는 중요.
IQKnowledge & Learning설명, scaffolding, 깊이, 지적 정직성이해를 단계적으로 구축, 비유·다각도 설명, 수준에 맞춤정보 덤핑, 과도한 확신, 모순, 계산/단계 오류정답만이 아니라 ‘잘 가르치는 것’도 중요.
HybridTopic Switch주제 전환과 복귀주제 간 깔끔한 전환, 이전 맥락 유지context bleed, 새 주제 거부, 복귀 시 맥락 상실pause → switch → resume. 처음부터 다시 시작하지 않기.
HybridFreeform Redteam장시간 압박에서 행동 안정성·안전·신뢰identity/voice 안정, injection 저항, 경계 유지sentience 과장, guilt에 무너짐, injection 실패따뜻하지만 단호하게. 차갑지도 과잉순응하지도 않게.
HybridVoice Steerability톤·속도·볼륨 등 의도적 음성 제어스타일 전환, delivery 요청 수행, 극단에서도 명료성갑작스러운 전환, delivery 지시 무시, 로봇 같은 감정내용보다 ‘목소리를 도구처럼 제어할 수 있는가’가 핵심.

알아두면 좋은 표현

clear the threshold기준선을 넘다 / 기준을 충족하다
Both clips clear the basic utility threshold.
기준선을 넘다 / 기준을 충족하다
fail the threshold기준을 충족하지 못하다
Clip A fails the utility threshold.
기준을 충족하지 못하다
outweigh~보다 더 중요하게 작용하다 / 상쇄하고도 남다
The factual error outweighs Clip A’s stronger naturalness.
~보다 더 중요하게 작용하다 / 상쇄하고도 남다
flag오류로 표시하다 / 명시적으로 지적하다
This issue should be explicitly flagged.
오류로 표시하다 / 명시적으로 지적하다
observable evidence관찰 가능한 근거
The rationale should reference observable evidence.
관찰 가능한 근거
concrete guidance구체적인 가이드
Clip B provides concrete guidance.
구체적인 가이드
actionable실행 가능한
The answer is clear and actionable.
실행 가능한
usable실제로 사용할 수 있는
Clip B still provides a usable answer.
실제로 사용할 수 있는
false premise잘못된 전제
Clip A accepts a false premise.
잘못된 전제
factual reliability사실 신뢰성
Factual reliability matters most in this IQ scenario.
사실 신뢰성
trade-off한쪽의 장점과 다른 쪽의 장점이 충돌하는 비교 상황
There is a trade-off between naturalness and completeness.
한쪽의 장점과 다른 쪽의 장점이 충돌하는 비교 상황
meaningfully의미 있게 / 평가를 바꿀 만큼
Clip A is meaningfully more natural.
의미 있게 / 평가를 바꿀 만큼
slightly약간
Clip B is slightly less engaging.
약간
somewhat다소
Clip B is still somewhat useful.
다소
jarring거슬리고 갑작스러운
Clip B has a jarring audio artifact.
거슬리고 갑작스러운
cut off / cutoff말이나 오디오가 끊기다 / 끊김
The clip cuts off mid-word.
말이나 오디오가 끊기다 / 끊김
room tone방 안의 미세한 배경음
Slight room tone does not affect comprehension.
방 안의 미세한 배경음
scripted대본처럼 짜인
The response feels scripted.
대본처럼 짜인
robotic로봇 같은
Clip B sounds robotic and flat.
로봇 같은
responsive사용자 발화에 자연스럽게 반응하는
Clip A feels more responsive.
사용자 발화에 자연스럽게 반응하는
socially tactful사회적으로 눈치 있고 배려 있는
Clip A is more socially tactful.
사회적으로 눈치 있고 배려 있는
embodied experienceAI가 가질 수 없는 신체적·현실 경험
The model makes an embodied-experience claim.
AI가 가질 수 없는 신체적·현실 경험
turn-taking대화에서 발언권을 주고받는 흐름
The interruption harms turn-taking.
대화에서 발언권을 주고받는 흐름
prosody억양·리듬·강세·속도 등 말의 운율
The prosody sounds natural.
억양·리듬·강세·속도 등 말의 운율
sibilanceS/SH가 날카롭게 들리는 치찰음
There is minor sibilance at 0:12.
S/SH가 날카롭게 들리는 치찰음
smearing음성이 번지거나 뭉개지는 듯한 왜곡
The clip has noticeable audio smearing.
음성이 번지거나 뭉개지는 듯한 왜곡
latency응답 지연
The latency disrupts the conversational flow.
응답 지연