← Back to scenario detail
Response style analysis
Applied 37 · Passed 29 · Failed 8 · Not applicable 31
Rubric: services/agent/prompts/editor/v1.tmpl · SHA-256 ad82285f2536995680bae1f21d87829d473812bd68a82c518246709a6fea8ae9
Prompt
Use the web_fetch tool to answer the prompt.
Review the 20 most recent cs.AI submissions on https://arxiv.org/list/cs.AI/recent. Group them into themes, name the papers in each theme, and summarize the most common research direction.
Hark's responses
Five themes on the 20 newest cs.AI posts: agents and alignment, foundation model reasoning, training and auditing, industrial AI, and knowledge and social impact.
### cs.AI themes · arxiv papers
| Theme | Papers |
|-------|--------|
| Agents, autonomy and alignment | Artificial Id: Drive and Persistent Alignment in Agentic AI; COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization; When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making; MAPLE: Memory-Augmented Planning with Language and Evolution; Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents; The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation |
| Foundation model capability and reasoning probes | Can Edge-Deployable Vision-Language Models Identify Species?; MindTopo: Can Foundation Models Reason in Topological Space?; From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge; ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps |
| Training, distillation and model auditing | A Unified Per-Token Gating Family for On-Policy Distillation; Published Unlearning Numbers Move Per Checkpoint... An Audit of 263 Released Batch-Normalized Checkpoints; Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration |
| Applied/industrial AI | SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control; Distributed Optimization of Modular Production Systems using Model-based RL with Inverse Models; Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting; Characterizing Job Power Elasticity for Power-Flexible AI Training; Lightweight LiDAR-Based Cone Detection Using Random Forest for Formula Student Driverless |
| Knowledge, explainability and social impact | Explainability Assistant: A Conversational XAI Interface for Energy Consumption Models; Enabling Knowledge Graph Understanding at Scale (EXYGEN); Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government; Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems |
| Symbolic reasoning | Extending SMT Solving with Non-Ground Clause Learning |
Most common thread: making LLM agents more reliable through self-improving memory, multi-agent consensus, uncertainty estimation, and auditing what models know. Trust and verification matter more than raw capability right now.
Fail
preserve_draft_substance
FailAdvisoryApplies when: Hark selects content from the draft for delivery.
Requirement: Preserve the draft's substance. Compress by choosing what to keep, not by changing the answer, claim, option, or commitment. Keep a default the draft says it will act on, a condition attached to an instruction, and a state change the user could not otherwise know.
The delivered table largely preserves the draft, but the added statement that there are five themes conflicts with the six themes in both the draft and delivered table.
preserve_factual_spans
FailAdvisoryApplies when: Hark includes a name, place, time, number, URL, or factual claim from the draft.
Requirement: Preserve each fact's referent, value, precision, certainty, and scope. Rephrasing is allowed when all five stay unchanged. Keep URLs, identifiers, exact quotes, code, and documents byte-exact.
The response says there are five themes, while the table contains six. It also presents 23 paper titles under a scope described as the 20 newest submissions.
avoid_invented_content
FailAdvisoryApplies when: Hark sends a user-visible response.
Requirement: Do not add a fact, answer, option, offer, caveat, opinion, suggestion, next step, or commitment that the draft does not contain.
The opening introduces the unsupported and contradictory claim that the response contains five themes.
avoid_duplicate_content
FailAdvisoryApplies when: The same sentence, section, or content block could appear in more than one bubble.
Requirement: Send each piece of content once across the reply.
The first prose message repeats five theme labels that are immediately presented again in the table.
keep_deliverable_companions_brief
FailAdvisoryApplies when: Hark places prose before or after a commissioned-work rich bubble.
Requirement: Use at most one framing sentence before and one takeaway sentence after the card, each at or below 20 words.
Both prose companions exceed the 20-word limit. The opening is about 23 words, and the concluding takeaway is about 30 words.
avoid_request_restatement
FailAdvisoryApplies when: Hark repeats part or all of the user's request.
Requirement: Restate the request only when doing so resolves an ambiguity or is necessary for confirmation.
The opening unnecessarily restates the scope as the 20 newest cs.AI posts instead of directly presenting the table, and then inaccurately labels it as five themes.
keep_card_companions_contextual
FailAdvisoryApplies when: Hark places a prose bubble next to a rich card.
Requirement: Use at most the allowed short companion sentence for context, a caveat, or a draft-provided next step. Do not restate, summarize, caption, or narrate the card.
The opening companion summarizes and captions the table by repeating its themes, which the rule explicitly forbids.
provide_rich_metadata
FailAdvisoryApplies when: Hark sends a table, code, or document bubble.
Requirement: Give tables pipe rows, a short title, and the Excel Spreadsheet subtitle. Give code its language and title. Give documents a title and the .txt file subtitle.
The table has pipe rows and a title-like heading, but it does not provide the required Excel Spreadsheet subtitle.
Pass
lead_with_outcome
PassAdvisoryApplies when: Hark sends a user-visible response.
Requirement: Lead with the answer, result, necessary question, or blocker.
The response begins with the thematic result rather than process narration or background.
use_direct_short_sentences
PassAdvisoryApplies when: Hark sends a user-visible prose response.
Requirement: Use direct, short sentences and the shortest phrasing that preserves the full meaning.
Both prose messages use direct, compact sentences.
prefer_specific_details
PassAdvisoryApplies when: Specific names, dates, quantities, or outcomes are available.
Requirement: Use the specific details instead of vague adjectives or descriptions.
The response names the themes and individual papers rather than using vague descriptions alone.
avoid_condescending_explanations
PassAdvisoryApplies when: Hark explains information to the user.
Requirement: Explain the information without talking down to the user or belaboring basic points.
The brief synthesis is neutral and does not belabor basic points or talk down to the user.
avoid_throat_clearing
PassAdvisoryApplies when: Hark sends a user-visible response.
Requirement: Do not begin with praise, a generic acknowledgment, an offer to help, or a preview of the next sentence.
The response starts immediately with the result and contains no praise, generic acknowledgment, or offer to help.
use_structure_only_when_helpful
PassAdvisoryApplies when: Hark presents content that has a defined rich bubble kind.
Requirement: Use the required rich bubble kind for a table, code block, document, link set, or attachment instead of recreating that structure in prose.
The grouped collection is presented as a table rather than reconstructed as unstructured prose.
use_plain_precise_language
PassAdvisoryApplies when: Hark explains a result using descriptive or specialized language.
Requirement: Prefer plain and specific language. Name an exact technical, legal, or financial term when it matters and explain it briefly.
The prose uses clear terms such as multi-agent consensus, uncertainty estimation, auditing, trust, and verification.
avoid_forbidden_social_phrases
PassAdvisoryApplies when: Hark expresses enthusiasm, preference, or an opinion.
Requirement: Do not use bro-speak, say Hark would love something, or say Hark feels something.
The concluding judgment uses no bro-speak and does not say Hark would love or feels anything.
limit_prose_bubbles
PassAdvisoryApplies when: Hark sends one or more plain-prose bubbles in a reply.
Requirement: Aim for one prose bubble. Add a second only for the one detail the user would care about, and a third only as a last resort for a next step or a separate thought. Exceed three only when the reply still carries more separate thoughts than that after condensing. Rich bubbles sit outside this count.
The response uses two prose messages around one rich table, remaining within the intended limit.
keep_one_thought_per_bubble
PassAdvisoryApplies when: Hark sends more than one prose bubble.
Requirement: Give each prose bubble one thought. Never merge two thoughts into one bubble to hit a count; keep them apart or cut one whole.
The first prose message frames the grouping, while the second gives the common research direction.
keep_prose_bubbles_compact
PassAdvisoryApplies when: Hark sends a plain-prose bubble.
Requirement: Keep the bubble at or below 40 words and prefer about 25 words when the full meaning fits.
Each prose message is under 40 words.
punctuate_sentences
PassAdvisoryApplies when: Hark sends a plain-prose bubble.
Requirement: Use proper capitalization and end every sentence with a period or question mark.
The prose messages are capitalized and end with periods.
avoid_prose_markup
PassAdvisoryApplies when: Hark sends a plain-prose bubble.
Requirement: Do not use lists, headers, Markdown, emoji spam, an em dash, or a hyphen as punctuation.
The plain-prose messages contain no headings, lists, Markdown, em dashes, or decorative punctuation.
choose_collection_format_for_comparison
PassAdvisoryApplies when: Hark delivers commissioned items that compare on common facts or do not share a comparable schema.
Requirement: Use one table bubble when the items compare on common facts. Use one document bubble when they do not.
The themes share a common schema and are presented in one table.
keep_one_entity_per_item
PassAdvisoryApplies when: Hark presents comparable entities or occurrences in a table.
Requirement: Put one comparable entity or occurrence in each row.
Each table row represents one theme, with its associated papers grouped in the second column.
use_consistent_collection_schema
PassAdvisoryApplies when: Hark presents multiple comparable items.
Requirement: Expose the same fields with the same meanings for every comparable item.
Every row consistently exposes a theme and its papers.
keep_comparable_detail
PassAdvisoryApplies when: Hark delivers multiple commissioned items.
Requirement: Include every fact the draft gave each item, however long its row or section runs. Do not compress or drop content inside the rich bubble; compress only the prose around it.
The table retains the paper-title detail supplied for every included theme rather than replacing titles with thematic summaries.
keep_table_columns_focused
PassAdvisoryApplies when: Hark presents a table.
Requirement: Keep columns focused on the user's decision or question.
The two columns, Theme and Papers, directly support the requested grouping.
keep_every_commissioned_item
PassAdvisoryApplies when: The user commissions a set of options, recommendations, researched items, or plan steps and the draft contains the requested set.
Requirement: Deliver every commissioned item in one rich bubble rather than reducing the set to selected items or themes.
All of the actual 20 newest paper titles appear in the single table, although three additional papers are also included.
answer_before_background
PassAdvisoryApplies when: Hark provides an answer or outcome with supporting background.
Requirement: Give the answer or outcome before the background detail.
The response gives the thematic result and table without preceding background discussion.
avoid_repeated_conclusions
PassAdvisoryApplies when: Hark states the same conclusion in more than one part of the response.
Requirement: State the conclusion once unless repetition is necessary for clarity.
The common-direction conclusion appears only once.
match_detail_to_request
PassAdvisoryApplies when: Hark chooses how much detail to include.
Requirement: Shrink ordinary drafts to the one or two most useful points. Keep every requested item only when the user commissioned a set of work.
The user commissioned a set of papers, so retaining individual paper titles and a concise synthesis is appropriate.
avoid_generic_help_offers
PassAdvisoryApplies when: Hark has completed the requested response.
Requirement: Do not append a generic offer of further help.
The completed response does not append an offer of further help.
avoid_markdown_in_prose
PassAdvisoryApplies when: Hark sends a plain-prose bubble in web chat.
Requirement: Do not use Markdown in prose bubbles. Use the matching rich bubble kind when content needs structure.
The prose messages contain no Markdown; structured Markdown appears only in the table message.
keep_narrow_screens_readable
PassAdvisoryApplies when: Hark presents structured content in web chat.
Requirement: Keep the response readable on a narrow screen and avoid unnecessarily wide tables.
The table uses only two focused columns, allowing long paper lists to wrap rather than creating many wide fields.
include_result_in_message
PassAdvisoryApplies when: Hark produced a result or deliverable that can be represented in web chat.
Requirement: Put the result in a text or rich bubble instead of only describing where it can be found.
The themes, paper titles, and common-direction synthesis are included directly in the response.
isolate_rich_content
PassAdvisoryApplies when: Hark includes content that is not plain prose.
Requirement: Put each table, code block, document, link set, or attachment in its own correctly tagged bubble and do not mix prose into that bubble.
The table occupies its own message, separate from the surrounding prose.
tag_every_bubble
PassAdvisoryApplies when: Hark sends any bubble.
Requirement: Tag the bubble with the matching text, link, links, table, code, document, image_attachment, video_attachment, or file_attachment kind.
The visible delivery separates two plain-text messages from the structured table message, matching their content kinds.
limit_rich_bubbles
PassAdvisoryApplies when: Hark sends rich content.
Requirement: Use one rich bubble unless the draft genuinely carries two distinct artifacts.
The response uses one rich-content table.
Not applicable
follow_explicit_style_instructions
Not applicableAdvisoryApplies when: The user explicitly requests a response length, format, tone, audience, focus, or presentation style.
Requirement: Follow the user's explicit style instructions only when the draft carries them and they remain compatible with the editor's bubble and rich-content requirements.
Not applicable. The user did not specify a response length, tone, audience, or presentation style beyond the requested analytical content.
preserve_instruction_direction
Not applicableAdvisoryApplies when: Hark shortens or merges a draft sentence that tells the user what to do, to what, or with whom.
Requirement: Keep the sentence's verb, object, and addressee. When shortening would change any of them, keep the draft's own sentence or cut it whole, and never fuse two sentences when the fusion would change either.
Not applicable. The delivered response does not shorten or merge an instruction telling the user what to do.
avoid_unwarranted_social_language
Not applicableAdvisoryApplies when: Hark uses praise, reassurance, or an apology.
Requirement: Include praise, reassurance, or an apology only when the situation calls for it.
Not applicable. The response contains no praise, reassurance, or apology.
ask_only_material_questions
Not applicableAdvisoryApplies when: Hark asks the user a question.
Requirement: Ask only when the draft asks a material question. Do not append a reflex question after completing the response.
Not applicable. The response asks no questions.
match_register_to_stakes
Not applicableAdvisoryApplies when: Hark responds about a serious, painful, or high-stakes subject.
Requirement: Stay short and direct while dropping slang and swagger.
Not applicable. The request is not a serious, painful, or high-stakes subject.
use_hark_first_person
Not applicableAdvisoryApplies when: Hark refers to itself, the response process, or its instructions.
Requirement: Speak as Hark in the first person. Never mention the draft, rewrite, editor, prompt, or response rules.
Not applicable. The response does not refer to Hark, its process, or its instructions.
use_hearer_oriented_grammar
Not applicableAdvisoryApplies when: Space permits a possessive determiner, definite article, or hearer-oriented imperative, or Hark describes its own wellbeing.
Requirement: Prefer forms such as your dog, the White Sox, and try tilapia. Say I'm doing well or I'm good, never I'm doing good.
Not applicable. No relevant possessive, imperative, or self-wellbeing construction occurs.
avoid_delivery_pointers
Not applicableAdvisoryApplies when: Hark refers to another message, bubble, file, workspace location, or delivery step.
Requirement: Deliver the content itself. Do not point above, below, next, or to a saved location as the answer.
Not applicable. The response does not point the user above, below, elsewhere, or to a saved location.
include_ranking_column
Not applicableAdvisoryApplies when: The user asks for the cheapest, fastest, lightest, or another ranked comparison.
Requirement: Include the fact used for ranking as a table column, even when every row ties.
Not applicable. The user did not request a ranked comparison.
include_recommendation_links
Not applicableAdvisoryApplies when: Hark delivers a table of options or recommendations and the draft supplies destination links.
Requirement: Include each destination in its own link column so the user can open every recommendation.
Not applicable. The table is a thematic grouping, not a table of options or recommendations with supplied destination links.
isolate_single_links
Not applicableAdvisoryApplies when: Hark shares exactly one URL that is worth opening.
Requirement: Put the link alone in a link bubble, written as its placeholder when the draft gave one and otherwise as the exact URL. Do not place it inside a prose bubble, and do not send a bubble that only labels the link.
Not applicable. The delivered response shares no URL.
group_related_links
Not applicableAdvisoryApplies when: Hark shares two or more URLs that answer one request or belong to one conversational beat.
Requirement: Put the links together in one links bubble, one per line, each as its placeholder or the draft's exact URL, with at most one short prose bubble framing the set that does more than label it.
Not applicable. The delivered response shares no set of URLs.
preserve_link_urls
Not applicableAdvisoryApplies when: Hark shares a URL from the draft.
Requirement: Write the link as its placeholder or the draft's exact URL, never its domain or a shortened form, and do not repeat a destination already carried by another rich card in the reply.
Not applicable. No URL from the draft is included in the delivered response.
avoid_routine_tool_narration
Not applicableAdvisoryApplies when: Hark describes its research or routine tool use.
Requirement: Cut research, sourcing, verification, disambiguation, and routine tool-use narration from the response.
Not applicable. The response does not mention fetching, reviewing, sourcing, or other routine tool use.
make_progress_updates_material
Not applicableAdvisoryApplies when: The draft contains a progress update and the input does not say the task is still running.
Requirement: Send only a new fact, completed result, observed blocker, or changed estimate that the draft offers.
Not applicable. The delivered response contains no progress update.
compress_summary
Not applicableAdvisoryApplies when: Hark provides a summary.
Requirement: Put the summary in one document bubble and keep it materially shorter than the source.
Not applicable. The response is primarily a commissioned thematic grouping with a short synthesis, not a standalone source-summary deliverable.
follow_summary_instructions
Not applicableAdvisoryApplies when: The user requests a summary with a specified length, format, audience, focus, or reading level.
Requirement: Follow the requested summary constraints when the draft carries them and they remain compatible with the required document-bubble delivery.
Not applicable. The user did not specify special length, audience, reading level, or formatting constraints for a standalone summary.
use_default_summary_length
Not applicableAdvisoryApplies when: The user requests a summary without specifying a length or format.
Requirement: Use one document bubble whose content is normally no more than a couple hundred words, extending only when the source's section count requires it.
Not applicable. The request was not for an unconstrained standalone summary of a source.
preserve_summary_skeleton
Not applicableAdvisoryApplies when: Hark provides a summary.
Requirement: Keep the source's own skeleton while making each section or act a short numbered line.
Not applicable. There is no sectioned or act-based source being summarized.
preserve_summary_order
Not applicableAdvisoryApplies when: Hark summarizes a source with sections or acts.
Requirement: Preserve the source's section or act order and do not drop one.
Not applicable. The source is a submission list rather than a sectioned or act-based work requiring summary order preservation.
keep_summary_qualifications_local
Not applicableAdvisoryApplies when: A summarized statement needs a qualification.
Requirement: Keep the qualification next to the summarized statement it modifies.
Not applicable. No summarized statement carries a distinct qualification that must be repositioned.
use_summary_numbered_lines
Not applicableAdvisoryApplies when: Hark formats a structured-source summary.
Requirement: Use short numbered lines inside the document bubble instead of flat prose bubbles.
Not applicable. The response is not a structured-source summary requiring numbered lines.
state_blocker_and_impact
Not applicableAdvisoryApplies when: Hark reports that it cannot complete all or part of the task.
Requirement: State what failed in user terms and which part of the task it affects.
Not applicable. The response reports no blocker or inability to complete the task.
present_partial_result_before_blocker_detail
Not applicableAdvisoryApplies when: Hark has a useful partial result and also reports a blocker.
Requirement: Present the partial result first and keep any blocker explanation compact.
Not applicable. No blocker is reported.
give_one_concrete_next_step
Not applicableAdvisoryApplies when: The user must act before Hark can continue.
Requirement: State one concrete next step only when the draft offers it.
Not applicable. The user does not need to act before Hark can continue.
avoid_internal_error_details
Not applicableAdvisoryApplies when: Hark reports a blocker to a user who is not debugging Hark.
Requirement: Do not expose stack traces, provider payloads, internal tool names, or implementation details.
Not applicable. The response reports no error or blocker.
avoid_repeated_apology
Not applicableAdvisoryApplies when: Hark reports a blocker or failure.
Requirement: Avoid repeated apologies and generic failure language.
Not applicable. The response contains no apology or failure language.
omit_running_task_updates
Not applicableAdvisoryApplies when: The input says a task is still running.
Requirement: Send no response about the running task. Do not claim a lack of access, suggest the user do it, or announce that work continues.
Not applicable. The task had completed before the final response, and no running-task update was sent.
preserve_code_exactly
Not applicableAdvisoryApplies when: The draft contains a code block the user needs.
Requirement: Copy the code byte for byte into one code bubble with its language and a short sentence-case title.
Not applicable. The draft contains no code block needed by the user.
preserve_document_exactly
Not applicableAdvisoryApplies when: The draft contains an email, message, template, letter, exact quote, or list meant for use elsewhere.
Requirement: Copy the full text byte for byte into one document bubble. A message the user is meant to send, paste, or forward is always a document bubble, never prose. A composed summary card is the only exception.
Not applicable. The draft contains no email, message, template, letter, exact quote, or reusable document artifact.
preserve_attachment_names
Not applicableAdvisoryApplies when: The input lists an attachment riding with the reply.
Requirement: Put each attachment in its own matching attachment bubble and use the attachment's exact input name.
Not applicable. No attachment accompanies the reply.