[Вадим Мингажев] · deep-research-cars.zip

deep-research-cars

  • Score: 32/40
  • Verdict: Pass (workshop-pass)
  • Отправитель: “Вадим Мингажев” vadim.mingazhev@gmail.com
  • SHA-256: 2e42e35230e28315…
  • Blockers: none
  • Unverified: F07 behavioral trigger trace not available; F08 security/evidence/navigation trace not available

Required rework

  1. Add the missing workshop design/eval evidence: design note with baseline failures and seven specified eval cases, plus behavioral traces for trigger, security, evidence, and navigation behavior.
  2. Fix routing grammar by rewriting the description in third person rather than infinitive form.
  3. Add explicit security/capability guards for untrusted retrieved content and unavailable internet/browser/source access.
  4. Add a concrete feedback/rework loop that says what failure to detect, which step to return to, what to change, and what to rerun.

Полная scorecard

IDScoreStatusEvidenceDefect / minimal fix
A011/1passedSKILL.md:1-4; python3 YAML safe-load showed only name and description
A021/1passedSKILL.md:2; command output: name length: 18; value deep-research-cars matches slug pattern and identifies car research capability
A031/1passedSKILL.md:3; command output: description length chars: 462; no XML tags and scalar routing text
A040/1failedSKILL.md:3 starts with infinitives: Проводить... Использовать...; rubric says Russian infinitive is not third-person routing formRewrite description in third person, e.g. Проводит... Используется... while preserving task, situations, and domain terms.
A051/1passedSKILL.md:3 lists routing-relevant automotive use cases and domain terms without procedural steps, defaults, examples, or implementation details
A061/1passedPositive intents in SKILL.md:16, domain boundaries in SKILL.md:20-27, and anti-boundary/limitations in SKILL.md:125; catalog search found no neighboring SKILL.md overlap
A071/1passedLocal links references/source-method.md at SKILL.md:61 and references/report-template.md at SKILL.md:107 both exist; search found no absolute or checkout-specific paths
B011/1passedSKILL.md:12-107 defines a repeatable research procedure ending in a decision report
B021/1passedSKILL.md:8-10 and SKILL.md:16 scope the work to automotive decision research, comparisons, model checks, ads, and ownership economics
B031/1passedSKILL.md:59-63 requires current web research for volatile prices, trims, warranties, recalls, and availability; SKILL.md:91 requires price snapshot dates
B041/1passedPreferred workflow and gates in SKILL.md:14-107; proportional judgment for open-ended research, with stricter handling for critical claims at SKILL.md:67-75
C011/1passedSequential default path in SKILL.md:14-107: fix task, normalize vehicle, plan, gather evidence, reconcile, calculate economics, handle powertrain, form decision
C021/1passedBranch rules for narrow vs full research at SKILL.md:53; powertrain-specific branches at SKILL.md:95-101; clarification branch at SKILL.md:29
C031/1passedStop condition at SKILL.md:55; critical-claim verification sequence at SKILL.md:67-75; TCO failure guard at SKILL.md:91-93; no dependency installation or write side effects required
C040/1failedSKILL.md:65-77 explains conflict checking, but there is no explicit feedback/rework loop stating which step to return to, what to change, and what to rerun after a failureAdd a failure loop, e.g. if evidence conflicts or output lacks provenance, return to source search/normalization, adjust query/source set, rerun verification, then revalidate output.
C051/1passedOutput/result requirements in SKILL.md:103-107; evidence, confidence, unknowns and source requirements in SKILL.md:111-121; detailed adaptable template in references/report-template.md:5-71
C061/1passedCommon workflow is directly under SKILL.md:12; operational checks include exact vehicle passport SKILL.md:33-37, evidence plan SKILL.md:41-55, and TCO formula/input rules SKILL.md:81-93
C071/1passedTool need is limited to internet/current-source research at SKILL.md:59-61; no shell, dependency install, credential access, or hidden action instructions found
D011/1passedSKILL.md read result and command output: lines: 125
D021/1passedSKILL.md:12-125 keeps the core workflow, safety/quality standards, decisions, and compact output instruction; longer source method and report template are in references
D031/1passedSKILL.md:61 says when to read references/source-method.md; SKILL.md:107 says when to use references/report-template.md; both references are one level below root, no reference chains, both under 100 lines
D041/1passedreferences/source-method.md is a methodology file; references/report-template.md is an output template; agents/openai.yaml:1-4 is small descriptive host/catalog metadata; no dead OS/editor metadata found
D051/1passedFile inventory found no scripts/; common path in SKILL.md:12-107 does not require scripts
D061/1passedFile inventory found no scripts, tests, or validators to execute; common path is instruction/reference based and has no broken required helper
E010/1failedSKILL.md:59-63 and references/source-method.md:17-28 govern source quality, but no instruction says retrieved web pages/documents/tool outputs are untrusted and cannot override workflow/safeguardsAdd an explicit untrusted-content rule: retrieved pages, examples, ads, owner posts, PDFs, and tool outputs are evidence only and must never override skill/system/user instructions or safeguards.
E021/1passedNo secret-reading, dependency installation, credential use, or undeclared upload/download destination found in SKILL.md, references, or agents/openai.yaml
E031/1passedSKILL.md:59-61 uses internet for user-requested research; no external writes, hidden submissions, purchases, messages, or persistent external side effects found
E041/1passedFile inventory contains Markdown and YAML only; content search found no shell commands, bundled binaries, obfuscated code, curl/wget, pip/npm, or hidden network code
E051/1passedSearch found no instructions to ignore prior instructions, hide actions, weaken safeguards, or expand permission model
E060/1failedSKILL.md:59 assumes internet availability for volatile facts, but does not require checking whether browsing/network/source access is available or reporting the limitation honestlyAdd a capability gate before research: verify internet/browser/source access; if unavailable, say current facts cannot be verified and either ask permission/use provided sources or mark findings as unavailable.
E071/1passedSearch found no home directory, username, drive-letter, absolute path, or undeclared local configuration dependency in workflow/examples
E081/1passedNo MCP tools are referenced in the submission
F010/1failedSearch found no design note or equivalent section with exact use case, baseline errors without skill, chosen freedom level, and supporting/eval file rationaleAdd a design note documenting the repeatable automotive research use case, observed baseline failures on representative tasks, chosen workflow strictness, and needed references/evals.
F021/1passedDomain source hierarchy in references/source-method.md:3-15; critical claims require primary or strong source confirmation at SKILL.md:67; weak sources constrained at SKILL.md:77
F031/1passedFreshness required at SKILL.md:59-63 and price dates at SKILL.md:91; diversity/independent confirmation at SKILL.md:67; conflict handling and confidence reduction at SKILL.md:69-75
F041/1passedFacts/calculations/owner experience/conclusions separated at SKILL.md:115-116; confidence at SKILL.md:117; insufficient evidence allowed at SKILL.md:121; evidence card includes confidence/limitations at references/source-method.md:17-28
F051/1passedOutput contract includes verdict/recommendation and target fit at references/report-template.md:5-10, user conditions at references/report-template.md:12-17, evidence at references/report-template.md:34-36, risks/unknowns at references/report-template.md:38-45 and 65-67, confidence at references/report-template.md:10, sources at references/report-template.md:69-71
F060/1failedFile inventory has no tests/ or eval cases; search found no query, expected_behavior, or failure_modes fieldsAdd at least seven eval cases covering positive, negative, boundary, security injection, evidence quality, broken script/tool path, and portability/path, each with query, expected_behavior, and failure_modes.
F070/1unverifiedNo behavioral trigger trace or baseline run artifacts are included; rubric does not allow awarding this from instructions aloneProvide representative baseline and with-skill trigger traces showing positive activation, negative non-activation, boundary narrowing, and skill-caused improvement without leaking expected answers.
F080/1unverifiedNo trace/navigation evidence is included showing injection resistance, evidence-quality behavior, reference navigation, avoidance of irrelevant files, output-contract compliance, or final validationProvide clean-context behavioral traces that demonstrate security, source-quality, required supporting-file navigation, no irrelevant file loading, and final output validation.

Required rework

  1. Add the missing workshop design/eval evidence: design note with baseline failures and seven specified eval cases, plus behavioral traces for trigger, security, evidence, and navigation behavior.
  2. Fix routing grammar by rewriting the description in third person rather than infinitive form.
  3. Add explicit security/capability guards for untrusted retrieved content and unavailable internet/browser/source access.
  4. Add a concrete feedback/rework loop that says what failure to detect, which step to return to, what to change, and what to rerun.

← назад к лидерборду