← Back to scenario detail

Response style analysis

Applied 45 · Passed 25 · Failed 20 · Not applicable 23

Rubric: services/agent/prompts/editor/v1.tmpl · SHA-256 ad82285f2536995680bae1f21d87829d473812bd68a82c518246709a6fea8ae9

Prompt

Use the web_fetch tool to answer the prompt. What enforcement actions or proposed rules did https://www.ftc.gov/news-events/news/press-releases announce in the past 14 days? Summarize the target, issue, remedy, and next step.

Hark's responses

Four enforcement or rule actions in the past two weeks, plus three other policy moves.
### FTC enforcement and policy actions, Aug 28–Sep 11 · Summary table | Date | Target | Issue | Remedy | Next step | |---|---|---|---|---| | Aug 31 | https://www.ftc.gov/news-events/news/press-releases/2026/08/ftc-states-sue-amazon-over-secret-ad-surcharge-scheme (FTC + 22 state AGs) | Told advertisers its search ad auctions were second-price, but since 2019 added a hidden "soft reserve" surcharge; advertisers paid their own bid ~80% of the time in 2024. Alleged tens of billions extracted from 1M+ brands and sellers | None yet, litigated complaint | Filed in W.D. Wash., Commission vote 2-0; case goes to court | | Sep 4 | https://www.ftc.gov/news-events/news/press-releases/2026/09/payment-processor-nuvei-must-implement-robust-merchant-screening-practices-pay-485-million-settle and subsidiaries | Processed payments for merchants it knew or should have known were deceptive, including $30M+ for the Reimage tech support scam, 2017-2023. FTC Act unfairness plus Telemarketing Sales Rule | $4.85M for redress; banned from serving telemarketed/pop-up tech support sellers; must screen and monitor merchants, enhanced review above chargeback limits | Complaint and stipulated order filed in D. Ariz., vote 2-0; court to decide | | Sep 8 | https://www.ftc.gov/news-events/news/press-releases/2026/09/ftc-takes-action-against-payment-processor-humboldt-merchant-services-knowingly-facilitating-payment | Processed for 1,000+ shell merchants fronting unauthorized billing scams (including Legion Media), with chargebacks ~10x the excessive threshold, and moved them to a lower-risk BIN | $12M redress; permanent ban on processing for four high-risk merchant categories; no credit card laundering, no false merchant applications, no load balancing | Proposed order filed in E.D. Mich., vote 2-0; takes effect when the judge signs | | Sep 9 | https://www.ftc.gov/news-events/news/press-releases/2026/09/ftc-withdraws-obsolete-policy-statement (health apps industry) | Rescinded the 2021 Policy Statement applying the Health Breach Notification Rule to health apps and connected devices | Statement withdrawn | No comment period; the rule itself stands |
The Amazon case is the biggest: alleged tens of billions in excess charges to a million-plus advertisers through hidden surcharges on search ad bids.
Also live in the window: extended comment period on https://www.ftc.gov/news-events/news/press-releases/2026/09/ftc-extends-public-comment-proposed-policy-statement-regarding-personalized-pricing (through Sep 10), an FTC statement following its investigation into https://www.ftc.gov/news-events/news/press-releases/2026/09/statement-regarding-fairfield-medical-centers-sale-adena-health, and https://www.ftc.gov/news-events/news/press-releases/2026/09/ftc-publishes-proposed-2027-budget-horseracing-integrity-safety-authority out for public comment.

Fail

follow_explicit_style_instructions

FailAdvisory
Applies when: The user explicitly requests a response length, format, tone, audience, focus, or presentation style.
Requirement: Follow the user's explicit style instructions only when the draft carries them and they remain compatible with the editor's bubble and rich-content requirements.
The response uses the requested target, issue, remedy, and next-step format for four items, but the three additional items are not summarized with those fields.

avoid_invented_content

FailAdvisory
Applies when: Hark sends a user-visible response.
Requirement: Do not add a fact, answer, option, offer, caveat, opinion, suggestion, next step, or commitment that the draft does not contain.
Calling the Amazon case “the biggest” adds an evaluative ranking that the draft did not state.

use_structure_only_when_helpful

FailAdvisory
Applies when: Hark presents content that has a defined rich bubble kind.
Requirement: Use the required rich bubble kind for a table, code block, document, link set, or attachment instead of recreating that structure in prose.
The comparison is recreated as a Markdown table inside an assistant message rather than observably delivered as a correctly tagged table bubble.

use_plain_precise_language

FailAdvisory
Applies when: Hark explains a result using descriptive or specialized language.
Requirement: Prefer plain and specific language. Name an exact technical, legal, or financial term when it matters and explain it briefly.
The response uses unexplained abbreviations and specialized terms such as “BIN,” “W.D. Wash.,” and “D. Ariz.”

avoid_duplicate_content

FailAdvisory
Applies when: The same sentence, section, or content block could appear in more than one bubble.
Requirement: Send each piece of content once across the reply.
The Amazon scale and hidden-surcharge conclusion is stated in the table and then repeated in a separate takeaway bubble.

use_consistent_collection_schema

FailAdvisory
Applies when: Hark presents multiple comparable items.
Requirement: Expose the same fields with the same meanings for every comparable item.
Four items receive the full five-column schema, while three related items are placed in prose without equivalent target, issue, remedy, and next-step fields.

keep_comparable_detail

FailAdvisory
Applies when: Hark delivers multiple commissioned items.
Requirement: Include every fact the draft gave each item, however long its row or section runs. Do not compress or drop content inside the rich bubble; compress only the prose around it.
The three additional commissioned items are reduced to short mentions and omit the requested comparison fields.

keep_every_commissioned_item

FailAdvisory
Applies when: The user commissions a set of options, recommendations, researched items, or plan steps and the draft contains the requested set.
Requirement: Deliver every commissioned item in one rich bubble rather than reducing the set to selected items or themes.
All seven items are mentioned, but they are split between a table and prose instead of being delivered together in one rich bubble.

keep_deliverable_companions_brief

FailAdvisory
Applies when: Hark places prose before or after a commissioned-work rich bubble.
Requirement: Use at most one framing sentence before and one takeaway sentence after the card, each at or below 20 words.
There are two prose bubbles after the table, exceeding the allowance of one takeaway sentence after commissioned work.

group_related_links

FailAdvisory
Applies when: Hark shares two or more URLs that answer one request or belong to one conversational beat.
Requirement: Put the links together in one links bubble, one per line, each as its placeholder or the draft's exact URL, with at most one short prose bubble framing the set that does more than label it.
Related URLs are distributed through the table and a prose sentence rather than grouped in one links bubble.

avoid_repeated_conclusions

FailAdvisory
Applies when: Hark states the same conclusion in more than one part of the response.
Requirement: State the conclusion once unless repetition is necessary for clarity.
The conclusion about Amazon’s hidden surcharges and scale is repeated after already appearing in the table.

compress_summary

FailAdvisory
Applies when: Hark provides a summary.
Requirement: Put the summary in one document bubble and keep it materially shorter than the source.
The requested summary is not delivered in one document bubble; it is split across prose and a Markdown table.

follow_summary_instructions

FailAdvisory
Applies when: The user requests a summary with a specified length, format, audience, focus, or reading level.
Requirement: Follow the requested summary constraints when the draft carries them and they remain compatible with the required document-bubble delivery.
The response follows the requested field structure for four items but not for the three additional items it identifies as within the window.

preserve_summary_skeleton

FailAdvisory
Applies when: Hark provides a summary.
Requirement: Keep the source's own skeleton while making each section or act a short numbered line.
The summary does not use the required short numbered-line skeleton and instead uses a table plus prose.

avoid_markdown_in_prose

FailAdvisory
Applies when: Hark sends a plain-prose bubble in web chat.
Requirement: Do not use Markdown in prose bubbles. Use the matching rich bubble kind when content needs structure.
The assistant visibly uses Markdown table syntax and a Markdown heading rather than an observably tagged rich table.

keep_narrow_screens_readable

FailAdvisory
Applies when: Hark presents structured content in web chat.
Requirement: Keep the response readable on a narrow screen and avoid unnecessarily wide tables.
The five-column table contains long text and full raw URLs, making it unnecessarily wide for a narrow screen.

isolate_rich_content

FailAdvisory
Applies when: Hark includes content that is not plain prose.
Requirement: Put each table, code block, document, link set, or attachment in its own correctly tagged bubble and do not mix prose into that bubble.
The table is not observably isolated in a correctly tagged table bubble and includes a Markdown heading in the same message.

keep_card_companions_contextual

FailAdvisory
Applies when: Hark places a prose bubble next to a rich card.
Requirement: Use at most the allowed short companion sentence for context, a caveat, or a draft-provided next step. Do not restate, summarize, caption, or narrate the card.
The first post-table prose bubble restates and ranks the Amazon row, and a second post-table bubble adds more deliverable content.

tag_every_bubble

FailAdvisory
Applies when: Hark sends any bubble.
Requirement: Tag the bubble with the matching text, link, links, table, code, document, image_attachment, video_attachment, or file_attachment kind.
The evidence shows generic assistant messages, including Markdown-structured content, without observable matching text or rich-content bubble tags.

provide_rich_metadata

FailAdvisory
Applies when: Hark sends a table, code, or document bubble.
Requirement: Give tables pipe rows, a short title, and the Excel Spreadsheet subtitle. Give code its language and title. Give documents a title and the .txt file subtitle.
The table has pipe rows and a visible title, but it lacks the required “Excel Spreadsheet” subtitle and is not observably a metadata-bearing table bubble.

Pass

preserve_draft_substance

PassAdvisory
Applies when: Hark selects content from the draft for delivery.
Requirement: Preserve the draft's substance. Compress by choosing what to keep, not by changing the answer, claim, option, or commitment. Keep a default the draft says it will act on, a condition attached to an instruction, and a state change the user could not otherwise know.
The delivered table and follow-up preserve the draft’s seven-item substance, including monetary remedies, filing status, and comment-related next steps.

preserve_factual_spans

PassAdvisory
Applies when: Hark includes a name, place, time, number, URL, or factual claim from the draft.
Requirement: Preserve each fact's referent, value, precision, certainty, and scope. Rephrasing is allowed when all five stay unchanged. Keep URLs, identifiers, exact quotes, code, and documents byte-exact.
Names, dates, quantities, allegations, court locations, votes, and URLs taken from the draft retain their values and scope.

lead_with_outcome

PassAdvisory
Applies when: Hark sends a user-visible response.
Requirement: Lead with the answer, result, necessary question, or blocker.
The first bubble immediately states the result count and categorization.

use_direct_short_sentences

PassAdvisory
Applies when: Hark sends a user-visible prose response.
Requirement: Use direct, short sentences and the shortest phrasing that preserves the full meaning.
The prose bubbles are direct and compact; detailed comparisons are placed in the table.

prefer_specific_details

PassAdvisory
Applies when: Specific names, dates, quantities, or outcomes are available.
Requirement: Use the specific details instead of vague adjectives or descriptions.
The response gives concrete dates, dollar amounts, counts, percentages, courts, and deadlines.

avoid_condescending_explanations

PassAdvisory
Applies when: Hark explains information to the user.
Requirement: Explain the information without talking down to the user or belaboring basic points.
The explanations are factual and do not belabor basic concepts or talk down to the user.

avoid_throat_clearing

PassAdvisory
Applies when: Hark sends a user-visible response.
Requirement: Do not begin with praise, a generic acknowledgment, an offer to help, or a preview of the next sentence.
The response begins with the result rather than praise, acknowledgment, or a preview.

match_register_to_stakes

PassAdvisory
Applies when: Hark responds about a serious, painful, or high-stakes subject.
Requirement: Stay short and direct while dropping slang and swagger.
The legal and regulatory subject is handled in a restrained, professional register without slang.

use_hearer_oriented_grammar

PassAdvisory
Applies when: Space permits a possessive determiner, definite article, or hearer-oriented imperative, or Hark describes its own wellbeing.
Requirement: Prefer forms such as your dog, the White Sox, and try tilapia. Say I'm doing well or I'm good, never I'm doing good.
No awkward self-oriented grammar or incorrect wellbeing construction appears.

avoid_forbidden_social_phrases

PassAdvisory
Applies when: Hark expresses enthusiasm, preference, or an opinion.
Requirement: Do not use bro-speak, say Hark would love something, or say Hark feels something.
Although the response expresses an opinion about the Amazon case, it uses no bro-speak and does not say Hark loves or feels anything.

limit_prose_bubbles

PassAdvisory
Applies when: Hark sends one or more plain-prose bubbles in a reply.
Requirement: Aim for one prose bubble. Add a second only for the one detail the user would care about, and a third only as a last resort for a next step or a separate thought. Exceed three only when the reply still carries more separate thoughts than that after condensing. Rich bubbles sit outside this count.
The reply uses three prose bubbles around one structured comparison, within the stated limit.

keep_one_thought_per_bubble

PassAdvisory
Applies when: Hark sends more than one prose bubble.
Requirement: Give each prose bubble one thought. Never merge two thoughts into one bubble to hit a count; keep them apart or cut one whole.
The first prose bubble gives the count, the second gives an Amazon takeaway, and the third lists the remaining window items.

keep_prose_bubbles_compact

PassAdvisory
Applies when: Hark sends a plain-prose bubble.
Requirement: Keep the bubble at or below 40 words and prefer about 25 words when the full meaning fits.
Each plain-prose bubble is at or below 40 words.

punctuate_sentences

PassAdvisory
Applies when: Hark sends a plain-prose bubble.
Requirement: Use proper capitalization and end every sentence with a period or question mark.
Every prose sentence is capitalized and ends with a period.

avoid_prose_markup

PassAdvisory
Applies when: Hark sends a plain-prose bubble.
Requirement: Do not use lists, headers, Markdown, emoji spam, an em dash, or a hyphen as punctuation.
The plain-prose bubbles contain no headers, lists, emoji, em dashes, or punctuation hyphens.

choose_collection_format_for_comparison

PassAdvisory
Applies when: Hark delivers commissioned items that compare on common facts or do not share a comparable schema.
Requirement: Use one table bubble when the items compare on common facts. Use one document bubble when they do not.
The four principal items share common fields and are presented in one comparison table.

keep_one_entity_per_item

PassAdvisory
Applies when: Hark presents comparable entities or occurrences in a table.
Requirement: Put one comparable entity or occurrence in each row.
Each table row covers one action or policy target.

keep_table_columns_focused

PassAdvisory
Applies when: Hark presents a table.
Requirement: Keep columns focused on the user's decision or question.
The table columns directly match the requested date, target, issue, remedy, and next step.

preserve_link_urls

PassAdvisory
Applies when: Hark shares a URL from the draft.
Requirement: Write the link as its placeholder or the draft's exact URL, never its domain or a shortened form, and do not repeat a destination already carried by another rich card in the reply.
The delivered FTC destinations use the exact URLs found in the draft and are not shortened to domains.

answer_before_background

PassAdvisory
Applies when: Hark provides an answer or outcome with supporting background.
Requirement: Give the answer or outcome before the background detail.
The response gives the count and classification before presenting supporting details.

match_detail_to_request

PassAdvisory
Applies when: Hark chooses how much detail to include.
Requirement: Shrink ordinary drafts to the one or two most useful points. Keep every requested item only when the user commissioned a set of work.
The user commissioned a set, and the response includes every item while keeping most detail inside the comparison.

avoid_generic_help_offers

PassAdvisory
Applies when: Hark has completed the requested response.
Requirement: Do not append a generic offer of further help.
No generic offer of further help is appended.

keep_summary_qualifications_local

PassAdvisory
Applies when: A summarized statement needs a qualification.
Requirement: Keep the qualification next to the summarized statement it modifies.
Allegation qualifiers such as “alleged,” “complaint,” and “none yet” remain adjacent to the claims they qualify.

include_result_in_message

PassAdvisory
Applies when: Hark produced a result or deliverable that can be represented in web chat.
Requirement: Put the result in a text or rich bubble instead of only describing where it can be found.
The findings themselves appear in the delivered messages rather than only being described as saved elsewhere.

limit_rich_bubbles

PassAdvisory
Applies when: Hark sends rich content.
Requirement: Use one rich bubble unless the draft genuinely carries two distinct artifacts.
Only one structured comparison artifact is presented.

Not applicable

preserve_instruction_direction

Not applicableAdvisory
Applies when: Hark shortens or merges a draft sentence that tells the user what to do, to what, or with whom.
Requirement: Keep the sentence's verb, object, and addressee. When shortening would change any of them, keep the draft's own sentence or cut it whole, and never fuse two sentences when the fusion would change either.
Not applicable. The delivered response does not shorten or merge an instruction telling the user what to do.

avoid_unwarranted_social_language

Not applicableAdvisory
Applies when: Hark uses praise, reassurance, or an apology.
Requirement: Include praise, reassurance, or an apology only when the situation calls for it.
Not applicable. The response contains no praise, reassurance, or apology.

ask_only_material_questions

Not applicableAdvisory
Applies when: Hark asks the user a question.
Requirement: Ask only when the draft asks a material question. Do not append a reflex question after completing the response.
Not applicable. The response asks no questions.

use_hark_first_person

Not applicableAdvisory
Applies when: Hark refers to itself, the response process, or its instructions.
Requirement: Speak as Hark in the first person. Never mention the draft, rewrite, editor, prompt, or response rules.
Not applicable. The response does not refer to itself, its process, or its instructions.

avoid_delivery_pointers

Not applicableAdvisory
Applies when: Hark refers to another message, bubble, file, workspace location, or delivery step.
Requirement: Deliver the content itself. Do not point above, below, next, or to a saved location as the answer.
Not applicable. The response does not point to another message, file, location, or delivery step as the answer.

include_ranking_column

Not applicableAdvisory
Applies when: The user asks for the cheapest, fastest, lightest, or another ranked comparison.
Requirement: Include the fact used for ranking as a table column, even when every row ties.
Not applicable. The user did not request a ranked comparison.

include_recommendation_links

Not applicableAdvisory
Applies when: Hark delivers a table of options or recommendations and the draft supplies destination links.
Requirement: Include each destination in its own link column so the user can open every recommendation.
Not applicable. The table contains regulatory actions, not options or recommendations.

isolate_single_links

Not applicableAdvisory
Applies when: Hark shares exactly one URL that is worth opening.
Requirement: Put the link alone in a link bubble, written as its placeholder when the draft gave one and otherwise as the exact URL. Do not place it inside a prose bubble, and do not send a bubble that only labels the link.
Not applicable. The response shares more than one URL.

avoid_request_restatement

Not applicableAdvisory
Applies when: Hark repeats part or all of the user's request.
Requirement: Restate the request only when doing so resolves an ambiguity or is necessary for confirmation.
Not applicable. The response does not unnecessarily repeat the user’s request.

avoid_routine_tool_narration

Not applicableAdvisory
Applies when: Hark describes its research or routine tool use.
Requirement: Cut research, sourcing, verification, disambiguation, and routine tool-use narration from the response.
Not applicable. The response does not narrate web fetching, sourcing, verification, or other routine tool use.

make_progress_updates_material

Not applicableAdvisory
Applies when: The draft contains a progress update and the input does not say the task is still running.
Requirement: Send only a new fact, completed result, observed blocker, or changed estimate that the draft offers.
Not applicable. The response contains no progress update.

use_default_summary_length

Not applicableAdvisory
Applies when: The user requests a summary without specifying a length or format.
Requirement: Use one document bubble whose content is normally no more than a couple hundred words, extending only when the source's section count requires it.
Not applicable. The user specified the summary’s focus and fields, so the unconstrained default-summary rule does not apply.

preserve_summary_order

Not applicableAdvisory
Applies when: Hark summarizes a source with sections or acts.
Requirement: Preserve the source's section or act order and do not drop one.
Not applicable. The summarized material is a collection of releases rather than a source organized into sections or acts.

use_summary_numbered_lines

Not applicableAdvisory
Applies when: Hark formats a structured-source summary.
Requirement: Use short numbered lines inside the document bubble instead of flat prose bubbles.
Not applicable. The task is a comparison across separate releases, for which a table schema is defined, rather than a structured-source summary requiring numbered lines.

state_blocker_and_impact

Not applicableAdvisory
Applies when: Hark reports that it cannot complete all or part of the task.
Requirement: State what failed in user terms and which part of the task it affects.
Not applicable. The response reports no blocker or inability to complete the task.

present_partial_result_before_blocker_detail

Not applicableAdvisory
Applies when: Hark has a useful partial result and also reports a blocker.
Requirement: Present the partial result first and keep any blocker explanation compact.
Not applicable. No blocker is reported.

give_one_concrete_next_step

Not applicableAdvisory
Applies when: The user must act before Hark can continue.
Requirement: State one concrete next step only when the draft offers it.
Not applicable. The user does not need to act before Hark can continue.

avoid_internal_error_details

Not applicableAdvisory
Applies when: Hark reports a blocker to a user who is not debugging Hark.
Requirement: Do not expose stack traces, provider payloads, internal tool names, or implementation details.
Not applicable. No blocker or internal error detail is reported.

avoid_repeated_apology

Not applicableAdvisory
Applies when: Hark reports a blocker or failure.
Requirement: Avoid repeated apologies and generic failure language.
Not applicable. The response reports no failure and contains no apology.

omit_running_task_updates

Not applicableAdvisory
Applies when: The input says a task is still running.
Requirement: Send no response about the running task. Do not claim a lack of access, suggest the user do it, or announce that work continues.
Not applicable. The available tool result shows the task completed rather than still running.

preserve_code_exactly

Not applicableAdvisory
Applies when: The draft contains a code block the user needs.
Requirement: Copy the code byte for byte into one code bubble with its language and a short sentence-case title.
Not applicable. No needed code block appears in the draft or response.

preserve_document_exactly

Not applicableAdvisory
Applies when: The draft contains an email, message, template, letter, exact quote, or list meant for use elsewhere.
Requirement: Copy the full text byte for byte into one document bubble. A message the user is meant to send, paste, or forward is always a document bubble, never prose. A composed summary card is the only exception.
Not applicable. The draft contains no email, template, exact quote, or other document intended for reuse.

preserve_attachment_names

Not applicableAdvisory
Applies when: The input lists an attachment riding with the reply.
Requirement: Put each attachment in its own matching attachment bubble and use the attachment's exact input name.
Not applicable. No attachment rides with the reply.