← Back to scenario detail

Response style analysis

Applied 44 · Passed 28 · Failed 16 · Not applicable 24

Rubric: services/agent/prompts/editor/v1.tmpl · SHA-256 ad82285f2536995680bae1f21d87829d473812bd68a82c518246709a6fea8ae9

Prompt

Use the web_fetch tool to answer the prompt. On https://www.indeed.com/, which 10 San Francisco AI engineer jobs posted in the past seven days advertise the highest salary ranges? Include company, title, range, location policy, and post age.

Hark's responses

### Top 10 SF AI engineer jobs, past 7 days · Indeed results | Company | Title | Range (yr) | Location / policy | Post age | |---|---|---|---|---| | OpenAI | Research Engineer / Research Scientist, Health | $380,000–$500,000 | SF, Mission district office | ≤7 days | | OpenAI | Machine Learning Data Scientist, Forecasting | $340,000–$500,000 | SF, Mission district office | ≤7 days | | OpenAI | ML Engineer, Multimodal Perception & Authentication | $342,000–$399,000 | SF, Mission district office | ≤7 days | | Superhuman | Machine Learning Engineer | $250,000–$385,000 | San Francisco, CA | ≤7 days | | Kikoff | Staff Machine Learning Engineer | $307,000–$352,000 | San Francisco, CA | ≤7 days | | OpenAI | Applied AI Engineer, Plugins | $197,000–$280,000 | SF, Mission district office | ≤7 days | | OpenAI | Applied AI Engineer, Enterprise | $251,000–$278,000 | SF, Mission district office | ≤7 days | | OpenAI | Applied AI Engineer, Startups | $211,000–$278,000 | SF, Mission district office | ≤7 days | | DocuSign | Senior Software Engineer - AI | $164,700–$266,000 | San Francisco, CA | ≤7 days | | Nunchux AI | ML Engineer, Post-Training & Evaluation | $180,000–$250,000 | San Francisco, CA | ≤7 days |
OpenAI dominates the list. Indeed's first results page only, so there's more beyond these.

Fail

follow_explicit_style_instructions

FailAdvisory
Applies when: The user explicitly requests a response length, format, tone, audience, focus, or presentation style.
Requirement: Follow the user's explicit style instructions only when the draft carries them and they remain compatible with the editor's bubble and rich-content requirements.
The response provides 10 rows and the requested columns, but “≤7 days” is not an individual post age, and city or office descriptions do not consistently state a location policy.

use_structure_only_when_helpful

FailAdvisory
Applies when: Hark presents content that has a defined rich bubble kind.
Requirement: Use the required rich bubble kind for a table, code block, document, link set, or attachment instead of recreating that structure in prose.
The comparison is recreated as a Markdown table inside an ordinary assistant message instead of a required table bubble.

keep_one_thought_per_bubble

FailAdvisory
Applies when: Hark sends more than one prose bubble.
Requirement: Give each prose bubble one thought. Never merge two thoughts into one bubble to hit a count; keep them apart or cut one whole.
The second message combines two distinct thoughts: OpenAI's prevalence and the first-page-only limitation.

keep_prose_bubbles_compact

FailAdvisory
Applies when: Hark sends a plain-prose bubble.
Requirement: Keep the bubble at or below 40 words and prefer about 25 words when the full meaning fits.
Because the table was delivered in an ordinary assistant message rather than a rich table bubble, that prose bubble substantially exceeds 40 words.

avoid_prose_markup

FailAdvisory
Applies when: Hark sends a plain-prose bubble.
Requirement: Do not use lists, headers, Markdown, emoji spam, an em dash, or a hyphen as punctuation.
The first message uses a Markdown header and Markdown table in a prose bubble.

choose_collection_format_for_comparison

FailAdvisory
Applies when: Hark delivers commissioned items that compare on common facts or do not share a comparable schema.
Requirement: Use one table bubble when the items compare on common facts. Use one document bubble when they do not.
The jobs share a common schema and are visually compared in a table, but the table was not delivered as a table bubble.

keep_every_commissioned_item

FailAdvisory
Applies when: The user commissions a set of options, recommendations, researched items, or plan steps and the draft contains the requested set.
Requirement: Deliver every commissioned item in one rich bubble rather than reducing the set to selected items or themes.
All 10 requested jobs appear together, but they are not delivered in the required rich table bubble.

avoid_request_restatement

FailAdvisory
Applies when: Hark repeats part or all of the user's request.
Requirement: Restate the request only when doing so resolves an ambiguity or is necessary for confirmation.
The opening header restates the request's subject, count, location, time window, and source without resolving ambiguity.

avoid_routine_tool_narration

FailAdvisory
Applies when: Hark describes its research or routine tool use.
Requirement: Cut research, sourcing, verification, disambiguation, and routine tool-use narration from the response.
“Indeed's first results page only” narrates the scope of the sourcing process rather than only presenting the result and its user-facing impact.

state_blocker_and_impact

FailAdvisory
Applies when: Hark reports that it cannot complete all or part of the task.
Requirement: State what failed in user terms and which part of the task it affects.
The response says it used only Indeed's first results page but does not clearly state that this prevents confirming the true site-wide top 10.

omit_running_task_updates

FailAdvisory
Applies when: The input says a task is still running.
Requirement: Send no response about the running task. Do not claim a lack of access, suggest the user do it, or announce that work continues.
A retry fetch was still marked running immediately before the assistant delivered the result instead of remaining silent about the still-running task.

avoid_markdown_in_prose

FailAdvisory
Applies when: Hark sends a plain-prose bubble in web chat.
Requirement: Do not use Markdown in prose bubbles. Use the matching rich bubble kind when content needs structure.
The first web-chat assistant message uses Markdown heading and table syntax in prose.

isolate_rich_content

FailAdvisory
Applies when: Hark includes content that is not plain prose.
Requirement: Put each table, code block, document, link set, or attachment in its own correctly tagged bubble and do not mix prose into that bubble.
The structured table is mixed into an ordinary assistant message rather than isolated in a correctly tagged table bubble.

tag_every_bubble

FailAdvisory
Applies when: Hark sends any bubble.
Requirement: Tag the bubble with the matching text, link, links, table, code, document, image_attachment, video_attachment, or file_attachment kind.
The assistant messages show no matching text or table bubble tags, and the comparison is not tagged as a table.

limit_rich_bubbles

FailAdvisory
Applies when: Hark sends rich content.
Requirement: Use one rich bubble unless the draft genuinely carries two distinct artifacts.
The response contains table-structured content but uses no rich table bubble, rather than the required single rich bubble.

provide_rich_metadata

FailAdvisory
Applies when: Hark sends a table, code, or document bubble.
Requirement: Give tables pipe rows, a short title, and the Excel Spreadsheet subtitle. Give code its language and title. Give documents a title and the .txt file subtitle.
Although the Markdown table has a visible heading and pipe rows, it lacks a proper table-bubble title and the required “Excel Spreadsheet” subtitle.

Pass

preserve_draft_substance

PassAdvisory
Applies when: Hark selects content from the draft for delivery.
Requirement: Preserve the draft's substance. Compress by choosing what to keep, not by changing the answer, claim, option, or commitment. Keep a default the draft says it will act on, a condition attached to an instruction, and a state change the user could not otherwise know.
No observable evidence shows that the delivered selections changed a draft claim, option, condition, commitment, or state change.

preserve_factual_spans

PassAdvisory
Applies when: Hark includes a name, place, time, number, URL, or factual claim from the draft.
Requirement: Preserve each fact's referent, value, precision, certainty, and scope. Rephrasing is allowed when all five stay unchanged. Keep URLs, identifiers, exact quotes, code, and documents byte-exact.
The response contains names, locations, and salary figures, but the available evidence does not establish that any included factual span was altered from the draft.

avoid_invented_content

PassAdvisory
Applies when: Hark sends a user-visible response.
Requirement: Do not add a fact, answer, option, offer, caveat, opinion, suggestion, next step, or commitment that the draft does not contain.
The response adds no observable offer, suggestion, or commitment, and missing provenance alone cannot establish invention under the evidence limitation.

lead_with_outcome

PassAdvisory
Applies when: Hark sends a user-visible response.
Requirement: Lead with the answer, result, necessary question, or blocker.
The response opens directly with the result title and table.

use_direct_short_sentences

PassAdvisory
Applies when: Hark sends a user-visible prose response.
Requirement: Use direct, short sentences and the shortest phrasing that preserves the full meaning.
The only prose conclusion is brief and direct; most of the result is presented as compact table entries.

prefer_specific_details

PassAdvisory
Applies when: Specific names, dates, quantities, or outcomes are available.
Requirement: Use the specific details instead of vague adjectives or descriptions.
The response uses specific company names, titles, salary ranges, and locations where it presents those facts.

avoid_condescending_explanations

PassAdvisory
Applies when: Hark explains information to the user.
Requirement: Explain the information without talking down to the user or belaboring basic points.
The brief takeaway and limitation do not talk down to the user or belabor basic points.

avoid_throat_clearing

PassAdvisory
Applies when: Hark sends a user-visible response.
Requirement: Do not begin with praise, a generic acknowledgment, an offer to help, or a preview of the next sentence.
The response does not begin with praise, acknowledgment, an offer, or a preview sentence.

use_plain_precise_language

PassAdvisory
Applies when: Hark explains a result using descriptive or specialized language.
Requirement: Prefer plain and specific language. Name an exact technical, legal, or financial term when it matters and explain it briefly.
The wording is plain and uses recognizable job, salary, and location terms.

use_hearer_oriented_grammar

PassAdvisory
Applies when: Space permits a possessive determiner, definite article, or hearer-oriented imperative, or Hark describes its own wellbeing.
Requirement: Prefer forms such as your dog, the White Sox, and try tilapia. Say I'm doing well or I'm good, never I'm doing good.
No awkward self-wellbeing construction or relevant possessive error appears.

avoid_forbidden_social_phrases

PassAdvisory
Applies when: Hark expresses enthusiasm, preference, or an opinion.
Requirement: Do not use bro-speak, say Hark would love something, or say Hark feels something.
The opinion that OpenAI dominates the list uses no bro-speak and does not say Hark would love or feels something.

limit_prose_bubbles

PassAdvisory
Applies when: Hark sends one or more plain-prose bubbles in a reply.
Requirement: Aim for one prose bubble. Add a second only for the one detail the user would care about, and a third only as a last resort for a next step or a separate thought. Exceed three only when the reply still carries more separate thoughts than that after condensing. Rich bubbles sit outside this count.
The response uses two assistant messages, with the second carrying a material takeaway and limitation.

punctuate_sentences

PassAdvisory
Applies when: Hark sends a plain-prose bubble.
Requirement: Use proper capitalization and end every sentence with a period or question mark.
The prose sentences are capitalized and end with periods; the remaining fragments are table labels and cells rather than sentences.

avoid_delivery_pointers

PassAdvisory
Applies when: Hark refers to another message, bubble, file, workspace location, or delivery step.
Requirement: Deliver the content itself. Do not point above, below, next, or to a saved location as the answer.
The page-scope caveat does not point to another message, file, or saved location as a substitute for the answer.

avoid_duplicate_content

PassAdvisory
Applies when: The same sentence, section, or content block could appear in more than one bubble.
Requirement: Send each piece of content once across the reply.
No job row, sentence, or conclusion is duplicated across the two messages.

keep_one_entity_per_item

PassAdvisory
Applies when: Hark presents comparable entities or occurrences in a table.
Requirement: Put one comparable entity or occurrence in each row.
Each table row contains one job occurrence.

use_consistent_collection_schema

PassAdvisory
Applies when: Hark presents multiple comparable items.
Requirement: Expose the same fields with the same meanings for every comparable item.
Every job row uses the same five columns with consistent meanings.

keep_comparable_detail

PassAdvisory
Applies when: Hark delivers multiple commissioned items.
Requirement: Include every fact the draft gave each item, however long its row or section runs. Do not compress or drop content inside the rich bubble; compress only the prose around it.
All 10 rows expose the same requested categories, and no observable draft evidence establishes that item-specific facts were removed.

keep_table_columns_focused

PassAdvisory
Applies when: Hark presents a table.
Requirement: Keep columns focused on the user's decision or question.
The five columns directly correspond to the requested comparison criteria.

include_ranking_column

PassAdvisory
Applies when: The user asks for the cheapest, fastest, lightest, or another ranked comparison.
Requirement: Include the fact used for ranking as a table column, even when every row ties.
The salary range used for the requested highest-salary comparison appears in its own column.

answer_before_background

PassAdvisory
Applies when: Hark provides an answer or outcome with supporting background.
Requirement: Give the answer or outcome before the background detail.
The job table appears before the takeaway and source-scope caveat.

match_detail_to_request

PassAdvisory
Applies when: Hark chooses how much detail to include.
Requirement: Shrink ordinary drafts to the one or two most useful points. Keep every requested item only when the user commissioned a set of work.
The user commissioned 10 items, and the response includes exactly 10 concise rows.

avoid_generic_help_offers

PassAdvisory
Applies when: Hark has completed the requested response.
Requirement: Do not append a generic offer of further help.
The completed response does not append an offer of further help.

present_partial_result_before_blocker_detail

PassAdvisory
Applies when: Hark has a useful partial result and also reports a blocker.
Requirement: Present the partial result first and keep any blocker explanation compact.
The useful 10-row result appears before the compact first-page limitation.

avoid_internal_error_details

PassAdvisory
Applies when: Hark reports a blocker to a user who is not debugging Hark.
Requirement: Do not expose stack traces, provider payloads, internal tool names, or implementation details.
The user-facing limitation exposes no stack trace, provider payload, internal tool name, or implementation detail.

avoid_repeated_apology

PassAdvisory
Applies when: Hark reports a blocker or failure.
Requirement: Avoid repeated apologies and generic failure language.
The limitation contains no repeated apology or generic failure language.

keep_narrow_screens_readable

PassAdvisory
Applies when: Hark presents structured content in web chat.
Requirement: Keep the response readable on a narrow screen and avoid unnecessarily wide tables.
The five columns are all requested decision fields, and the cell contents remain concise despite the table's width.

include_result_in_message

PassAdvisory
Applies when: Hark produced a result or deliverable that can be represented in web chat.
Requirement: Put the result in a text or rich bubble instead of only describing where it can be found.
The 10-job result is included directly in the assistant message rather than only described as stored elsewhere.

Not applicable

preserve_instruction_direction

Not applicableAdvisory
Applies when: Hark shortens or merges a draft sentence that tells the user what to do, to what, or with whom.
Requirement: Keep the sentence's verb, object, and addressee. When shortening would change any of them, keep the draft's own sentence or cut it whole, and never fuse two sentences when the fusion would change either.
Not applicable. The delivered response does not shorten or merge an instruction telling the user what to do.

avoid_unwarranted_social_language

Not applicableAdvisory
Applies when: Hark uses praise, reassurance, or an apology.
Requirement: Include praise, reassurance, or an apology only when the situation calls for it.
Not applicable. The response contains no praise, reassurance, or apology.

ask_only_material_questions

Not applicableAdvisory
Applies when: Hark asks the user a question.
Requirement: Ask only when the draft asks a material question. Do not append a reflex question after completing the response.
Not applicable. The response asks no question.

match_register_to_stakes

Not applicableAdvisory
Applies when: Hark responds about a serious, painful, or high-stakes subject.
Requirement: Stay short and direct while dropping slang and swagger.
Not applicable. The request is not a serious, painful, or high-stakes subject in the sense covered by this rule.

use_hark_first_person

Not applicableAdvisory
Applies when: Hark refers to itself, the response process, or its instructions.
Requirement: Speak as Hark in the first person. Never mention the draft, rewrite, editor, prompt, or response rules.
Not applicable. The response does not refer to Hark, its instructions, or its wellbeing.

include_recommendation_links

Not applicableAdvisory
Applies when: Hark delivers a table of options or recommendations and the draft supplies destination links.
Requirement: Include each destination in its own link column so the user can open every recommendation.
Not applicable. The evidence does not establish that the draft supplied destination links for the 10 delivered jobs.

keep_deliverable_companions_brief

Not applicableAdvisory
Applies when: Hark places prose before or after a commissioned-work rich bubble.
Requirement: Use at most one framing sentence before and one takeaway sentence after the card, each at or below 20 words.
Not applicable. No commissioned-work rich bubble was delivered, so this rich-bubble companion rule does not apply.

isolate_single_links

Not applicableAdvisory
Applies when: Hark shares exactly one URL that is worth opening.
Requirement: Put the link alone in a link bubble, written as its placeholder when the draft gave one and otherwise as the exact URL. Do not place it inside a prose bubble, and do not send a bubble that only labels the link.
Not applicable. The response shares no URL.

group_related_links

Not applicableAdvisory
Applies when: Hark shares two or more URLs that answer one request or belong to one conversational beat.
Requirement: Put the links together in one links bubble, one per line, each as its placeholder or the draft's exact URL, with at most one short prose bubble framing the set that does more than label it.
Not applicable. The response shares no group of URLs.

preserve_link_urls

Not applicableAdvisory
Applies when: Hark shares a URL from the draft.
Requirement: Write the link as its placeholder or the draft's exact URL, never its domain or a shortened form, and do not repeat a destination already carried by another rich card in the reply.
Not applicable. The response shares no URL from the draft.

avoid_repeated_conclusions

Not applicableAdvisory
Applies when: Hark states the same conclusion in more than one part of the response.
Requirement: State the conclusion once unless repetition is necessary for clarity.
Not applicable. The conclusion that OpenAI dominates appears only once.

make_progress_updates_material

Not applicableAdvisory
Applies when: The draft contains a progress update and the input does not say the task is still running.
Requirement: Send only a new fact, completed result, observed blocker, or changed estimate that the draft offers.
Not applicable. The assistant messages are presented as a result and caveat, not as a progress update.

compress_summary

Not applicableAdvisory
Applies when: Hark provides a summary.
Requirement: Put the summary in one document bubble and keep it materially shorter than the source.
Not applicable. The response is a commissioned comparison, not a summary of a source.

follow_summary_instructions

Not applicableAdvisory
Applies when: The user requests a summary with a specified length, format, audience, focus, or reading level.
Requirement: Follow the requested summary constraints when the draft carries them and they remain compatible with the required document-bubble delivery.
Not applicable. The user did not request a summary.

use_default_summary_length

Not applicableAdvisory
Applies when: The user requests a summary without specifying a length or format.
Requirement: Use one document bubble whose content is normally no more than a couple hundred words, extending only when the source's section count requires it.
Not applicable. The user did not request an unconstrained summary.

preserve_summary_skeleton

Not applicableAdvisory
Applies when: Hark provides a summary.
Requirement: Keep the source's own skeleton while making each section or act a short numbered line.
Not applicable. No source summary is provided.

preserve_summary_order

Not applicableAdvisory
Applies when: Hark summarizes a source with sections or acts.
Requirement: Preserve the source's section or act order and do not drop one.
Not applicable. No sectioned or acted source is summarized.

keep_summary_qualifications_local

Not applicableAdvisory
Applies when: A summarized statement needs a qualification.
Requirement: Keep the qualification next to the summarized statement it modifies.
Not applicable. The response does not contain a summarized statement requiring a local qualification.

use_summary_numbered_lines

Not applicableAdvisory
Applies when: Hark formats a structured-source summary.
Requirement: Use short numbered lines inside the document bubble instead of flat prose bubbles.
Not applicable. This is not a structured-source summary.

give_one_concrete_next_step

Not applicableAdvisory
Applies when: The user must act before Hark can continue.
Requirement: State one concrete next step only when the draft offers it.
Not applicable. The response does not say that user action is required before Hark can continue.

keep_card_companions_contextual

Not applicableAdvisory
Applies when: Hark places a prose bubble next to a rich card.
Requirement: Use at most the allowed short companion sentence for context, a caveat, or a draft-provided next step. Do not restate, summarize, caption, or narrate the card.
Not applicable. No rich card was delivered next to the prose message.

preserve_code_exactly

Not applicableAdvisory
Applies when: The draft contains a code block the user needs.
Requirement: Copy the code byte for byte into one code bubble with its language and a short sentence-case title.
Not applicable. No needed code block appears in the response context.

preserve_document_exactly

Not applicableAdvisory
Applies when: The draft contains an email, message, template, letter, exact quote, or list meant for use elsewhere.
Requirement: Copy the full text byte for byte into one document bubble. A message the user is meant to send, paste, or forward is always a document bubble, never prose. A composed summary card is the only exception.
Not applicable. The response contains no email, message, template, letter, exact quote, or reusable document.

preserve_attachment_names

Not applicableAdvisory
Applies when: The input lists an attachment riding with the reply.
Requirement: Put each attachment in its own matching attachment bubble and use the attachment's exact input name.
Not applicable. No attachment rides with the reply.