Skip to content

fix: find tool parameters inside composed schemas - #164032

Merged
obviyus merged 1 commit into
openclaw:mainfrom
harshitgupta31415:harshitgupta31415/tool-search-misses-composed-schema-parameters
Oct 3, 2026
Merged

obviyus merged 1 commit into
openclaw:mainfrom
harshitgupta31415:harshitgupta31415/tool-search-misses-composed-schema-parameters

Conversation

@harshitgupta31415

@harshitgupta31415 harshitgupta31415 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #164024, reported and reproduced by @harshitgupta31415.

What Problem This Solves

Tool Search misses a tool when the query appears only in parameter names or descriptions inside anyOf, oneOf, or allOf schemas.

User Impact

Tools with composed schemas become discoverable through their parameter metadata in both structured search and directory mode. Additional indexed words can affect relative search scores; exact-name priority remains covered by the existing tests.

Why This Change Was Made

The shared metadata walker followed object properties and array items but omitted composition branches. This extends that existing traversal, preserving its depth limit, trusted-source filtering, and cache refresh behavior. The same loop also handles tuple items. Production change: 5 additions, 3 deletions (+2 lines); the added child-schema edges require no separate index, configuration, or public API change.

Evidence

  • Reproduced on upstream 4a0be71e57ab19d08518879dcab11c5e9c798cd6: five added cases failed while all 33 existing ranking tests passed. The miss was expected [] to deeply equal ['indexed_resource'].
  • Candidate: 3a6e3673c37a520e6b03be24368ddf5b9c577f51.
  • After repair: all 41 ranking tests passed. Coverage includes each composition keyword, multiple branches, nested object/array schemas, tuple items, metadata edits, visibility filtering, cyclic depth limits, boolean schemas, literal-data exclusion, and both untrusted sources. Existing exact-name/ID precedence and lexical ranking cases passed.
  • Actual control-flow proof uses createToolSearchTools and the real catalog compaction, search, describe, and call paths. No search/runtime functions are mocked, and discovery never invokes the target tool. The same script failed before the fix and passed after it:
Mode Schema Parameter-name query Description query Describe / valid call Invalid / revoked call
tools anyOf found found passed denied
tools oneOf found found passed denied
tools allOf found found passed denied
directory anyOf found found passed denied
directory oneOf found found passed denied
directory allOf found found passed denied
  • Formatting: oxfmt 0.68.0 passed for both changed files on Linux. The Windows native binding was blocked by Application Control; security settings were not changed.
  • git diff --check passed. The complete diff was reviewed locally; only the production owner and its regression tests are included. Test delta: +83 lines, including expanding the existing untrusted-source case.

The broader Tool Search run passed all 181 tests in four files (273.20s total); core production typecheck and focused lint completed with exit 0 and no diagnostics. The committed-head runtime proof returned verdict: pass, six passing mode/keyword combinations, and an empty source diff.

The full changed-file gate stopped at the Windows SQLite worker ratchet (1339 -> 2083). Re-running that check with the entire tracked working tree restored to clean base 4a0be71e57ab19d08518879dcab11c5e9c798cd6 produced byte-identical diagnostic lines and exit 1; the candidate was then restored and verified clean. This baseline failure is unrelated to the metadata repair, and later checks in that gate were not executed. Conflict markers, line-cap, max-lines, and assertion-safety checks passed. Follow-up: investigate the Windows SQLite worker ratchet separately.

Changed-test typecheck (agents-tools) passed in 220.8s. CI run 37096762155 completed successfully on head 3a6e3673c37a520e6b03be24368ddf5b9c577f51; all selected test, lint, typecheck, guard, and Gateway checks passed, and openclaw/ci-gate is green. ClawSweeper review confirmed sufficient behavior proof and no actionable findings on the same head. Ready for maintainer review. The full repository build/test/docs suites have not been run. The original contributor proof was local discovery only; the maintainer Gateway and real-provider checks below extend that evidence.

Commands and reproducible runtime proof
node scripts/run-vitest.mjs src/agents/tool-search-ranking.test.ts --maxWorkers=1
node scripts/run-vitest.mjs src/agents/tool-search-ranking.test.ts src/agents/tool-search.test.ts src/agents/tool-search-runtime.test.ts src/agents/tool-search-config.test.ts --maxWorkers=1
node scripts/run-tsgo.mjs -p tsconfig.core.json --incremental --tsBuildInfoFile .artifacts/tsgo-cache/core.tsbuildinfo
node scripts/run-tsgo-core-test-shards.mjs --changed-paths-json '["src/agents/tool-search-ranking.test.ts"]'
node scripts/check-changed.mjs --base 4a0be71e57ab19d08518879dcab11c5e9c798cd6 --timed -- src/agents/tool-search-ranking.ts src/agents/tool-search-ranking.test.ts

Save the following script outside the checkout and run it from the repository root with node --import ./scripts/tsx.mjs /absolute/path/behavior-proof.mjs. It exits 1 on the base and 0 with the repair.

import assert from "node:assert/strict";
import { execFileSync } from "node:child_process";
import { resolve } from "node:path";
import { pathToFileURL } from "node:url";

const root = process.cwd();
const {
  applyToolSearchCatalog,
  applyToolSchemaDirectoryCatalog,
  createToolSearchCatalogRef,
  createToolSearchTools,
  restrictToolSearchCatalog,
} = await import(pathToFileURL(resolve(root, "src/agents/tool-search.ts")));
const observations = [];
const failures = [];
for (const mode of ["tools", "directory"]) {
  for (const keyword of ["anyOf", "oneOf", "allOf"]) {
    const parameters = {
      [keyword]: [{
        type: "object",
        properties: { orchard: { type: "string", description: "Collect apples" } },
        required: ["orchard"],
        additionalProperties: false,
      }],
    };
    let calls = 0;
    const tool = {
      name: "indexed_resource",
      label: "Indexed resource",
      description: "Process a resource",
      parameters,
      execute: async (_id, input) => {
        calls++;
        return {
          content: [{ type: "text", text: JSON.stringify(input) }],
          details: input,
        };
      },
    };
    const catalogRef = createToolSearchCatalogRef();
    const ctx = {
      config: { tools: { toolSearch: { enabled: true, mode } } },
      catalogRef,
      sessionId: `schema-proof-${mode}-${keyword}`,
    };
    const apply = mode === "directory" ? applyToolSchemaDirectoryCatalog : applyToolSearchCatalog;
    const result = apply({ ...ctx, tools: [...createToolSearchTools(ctx), tool] });
    assert.equal(result.compacted, true);
    const controls = new Map(result.tools.map((entry) => [entry.name, entry]));
    const execute = (name, input) => controls.get(name).execute(`proof-${name}`, input);
    const names = (result) => result.details.map((entry) => entry.name);
    const byName = names(await execute("tool_search", { query: "orchard" }));
    const byDescription = names(await execute("tool_search", { query: "apples" }));
    assert.equal(calls, 0, "discovery must not execute the target");
    for (const [query, found] of [["orchard", byName], ["apples", byDescription]]) {
      if (JSON.stringify(found) !== JSON.stringify([tool.name])) {
        failures.push({ mode, keyword, query, expected: [tool.name], actual: found });
      }
    }
    const described = await execute("tool_describe", { id: tool.name });
    assert.deepEqual(described.details.parameters, parameters);
    const called = await execute("tool_call", { id: tool.name, args: { orchard: "north" } });
    assert.deepEqual(called.details.result.details, { orchard: "north" });
    assert.equal(calls, 1);
    await assert.rejects(execute("tool_call", { id: tool.name, args: { orchard: 4 } }));
    assert.equal(calls, 1, "invalid input must not execute");
    restrictToolSearchCatalog({ catalogRef, allowedToolNames: new Set() });
    assert.deepEqual(names(await execute("tool_search", { query: "orchard" })), []);
    await assert.rejects(execute("tool_call", { id: tool.name, args: { orchard: "north" } }));
    observations.push({ mode, keyword, byName, byDescription, describe: "pass", validCall: "pass", invalidCall: "denied", revokedSearch: "hidden", revokedCall: "denied" });
  }
}
console.log(JSON.stringify({
  head: execFileSync("git", ["rev-parse", "HEAD"], { encoding: "utf8" }).trim(),
  sourceChanges: execFileSync("git", ["diff", "--stat"], { encoding: "utf8" }).trim(),
  node: process.version,
  verdict: failures.length ? "fail" : "pass",
  observations,
  failures,
}, null, 2));
process.exitCode = failures.length ? 1 : 0;

Maintainer Gateway and live verification

Verified the actual built Gateway path on current main df883d4cc7bdbfc724cc167ccb2b2e2ecfd727cd and candidate 3a6e3673c37a520e6b03be24368ddf5b9c577f51, in isolated Linux containers with Tool Search enabled.

An explicitly allowed local plugin registers indexed_resource with payload.anyOf containing an object property named orchard, described as Collect apples. Neither word appears in the tool name or description. This is an OpenClaw plugin catalog entry, within the documented trusted-schema boundary.

Gateway tool_search query Current main Candidate
orchard [] indexed_resource
apples [] indexed_resource
saffron (flat-schema control) flat_resource, flat_secondary Same matches and order
xylophone (unused $defs behind recursive $ref) [], completes [], completes

The first root-level single-branch union fixture was normalized before cataloging and therefore already matched on main; the nested union above preserves the composition at the real indexing boundary and discriminates the fix.

The candidate regression file run against current-main production code had 7 failures and 34 passes (22.22s); the candidate passed all 41 tests (23.10s). This includes all three composition keywords, nested schemas, unchanged ranking cases, both untrusted sources remaining untraversed, and bounded cyclic traversal.

Live check on the candidate: two real gpt-5-mini requests through a passive proxy, both HTTP 200, reasoning_effort: low, max_completion_tokens: 4096. The model called tool_search for orchard, received indexed_resource, and reported the discovered tool without executing it. Utility jobs were off. An earlier proxy preflight rejected requests missing the explicit reasoning field before any provider call; configuring the model's supported reasoning capability resolved that fixture setup issue. No request/response payloads were rewritten.

No overlap with Pash/Sarah changes. No source changes were needed on top of @harshitgupta31415's fix, and the guarded push confirmed the existing head. No CI jobs were rerun. CI and ClawSweeper are ready on that exact SHA; the PR is mergeable.

@clawsweeper

clawsweeper Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@openclaw-barnacle openclaw-barnacle Bot added agents Agent runtime and tooling size: S labels Oct 3, 2026
@harshitgupta31415
harshitgupta31415 marked this pull request as ready for review October 3, 2026 04:30
@clawsweeper clawsweeper Bot added P2 Normal backlog priority with limited blast radius. proof: sufficient ClawSweeper judged the real behavior proof convincing. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. labels Oct 3, 2026
@clawsweeper

clawsweeper Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Codex review: needs maintainer review before merge. Reviewed October 3, 2026, 12:47 AM ET / 04:47 UTC (Revision 2).

ClawSweeper review

What this changes

Tool Search now finds trusted tools through parameter names and descriptions inside composed schemas and tuple items.

Merge readiness

✅ Ready for maintainer review

This repair remains necessary on current main and the latest release. The pinned patch has sufficient behavior proof and no actionable correctness findings.

Priority: P2
Reviewed head: 3a6e3673c37a520e6b03be24368ddf5b9c577f51

Review scores

Measure Result What it means
Overall readiness 🐚 platinum hermit (4/6) A focused repair with production-path before/after evidence, useful regressions, and no blocking findings.
Proof confidence 🐚 platinum hermit (4/6) Sufficient (live_output): The supplied committed-head before/after results exercise actual catalog compaction and search controls in both modes across all three composition keywords, showing restored metadata discovery and preserved call rejection. No stored-data contract changes.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Verified Sufficient (live_output): The supplied committed-head before/after results exercise actual catalog compaction and search controls in both modes across all three composition keywords, showing restored metadata discovery and preserved call rejection. No stored-data contract changes.
Evidence reviewed 9 items Pinned introduced change: The introduced production patch adds composition and tuple traversal to the existing collector without changing its depth guard. The verified test-merge comparison contains the same production change.
Current-main necessity: Fetched main still traverses only properties and single-schema items, so it cannot index the reported orchard/apples metadata inside composition branches.
Latest-release comparison: The v2026.9.8 source also retains the properties/items-only walker. Local release-blob inspection failed with a lazy-fetch HTTP 403; the GitHub contents endpoint supplied the release source successfully.
Findings None None.
Security None None.

How this fits together

OpenClaw Tool Search turns the current policy-filtered tool catalog into searchable metadata. Its shared ranking index feeds discovery in both structured and directory modes before normal tool execution.

flowchart TD
 A[Available tool catalog] --> B[Visibility filtering]
 B --> C[Trusted schema metadata]
 C --> D[Bounded schema traversal]
 D --> E[Lexical search index]
 F[Search query] --> E
 E --> G[Ranked tool descriptors]
Loading

Before merge

None.

Agent review details

Security

None.

PR surface

Source +2, Tests +83. Total +85 across 2 files.

View PR surface stats
Area Files Added Removed Net
Source 1 5 3 +2
Tests 1 109 26 +83
Docs 0 0 0 0
Config 0 0 0 0
Generated 0 0 0 0
Other 0 0 0 0
Total 2 114 29 +85

Review metrics

Metric Value Why it matters
Production versus regression growth production net +2 lines; tests net +83 lines The small production increase extends the existing traversal; regression growth covers discovery and retained safeguards.

Root-cause cluster

Relationship: fixed_by_candidate
Canonical: #164024
Summary: This PR is the implementation candidate for the linked composed-schema discovery defect.

Members:

Proposal only: this assessment does not dispatch repair, suppress jobs, mutate sibling items, close, or merge anything.

Technical review

Best possible solution:

Keep one bounded metadata walker serving both discovery modes while preserving trusted-source filtering and existing execution contracts.

Do we have a high-confidence way to reproduce the issue?

Yes: current-main source deterministically omits orchard/apples metadata inside composition branches, and the contributor records before/after runs through both real control modes. This reviewer did not execute the reproduction.

Is this the best way to solve the issue?

Yes: extending the existing shared traversal repairs the documented discovery contract without another index, configuration option, or public API.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning medium; reviewed against 350e91cf7656.

Labels

Label changes:

No label changes.

Label justifications:

  • P2: This repairs a bounded discovery failure with exact-name lookup still available.
  • rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🐚 platinum hermit and patch quality is 🐚 platinum hermit.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Sufficient (live_output): The supplied committed-head before/after results exercise actual catalog compaction and search controls in both modes across all three composition keywords, showing restored metadata discovery and preserved call rejection. No stored-data contract changes.
  • proof: sufficient: Contributor real behavior proof is sufficient. The supplied committed-head before/after results exercise actual catalog compaction and search controls in both modes across all three composition keywords, showing restored metadata discovery and preserved call rejection. No stored-data contract changes.

Evidence

What I checked:

  • Pinned introduced change: The introduced production patch adds composition and tuple traversal to the existing collector without changing its depth guard. The verified test-merge comparison contains the same production change. (src/agents/tool-search-ranking.ts:28, 3a6e3673c37a)
  • Current-main necessity: Fetched main still traverses only properties and single-schema items, so it cannot index the reported orchard/apples metadata inside composition branches. (src/agents/tool-search-ranking.ts:28, 350e91cf7656)
  • Latest-release comparison: The v2026.9.8 source also retains the properties/items-only walker. Local release-blob inspection failed with a lazy-fetch HTTP 403; the GitHub contents endpoint supplied the release source successfully. (src/agents/tool-search-ranking.ts:28, fc23bc864e45)
  • Runtime and documented boundary: The runtime indexes parameters only for openclaw entries, compares collected text when refreshing cached indexes, and applies visibility before ranking. Documentation promises first-party parameter metadata search and explicitly separates these controls from Codex-native Tool Search; no Codex dependency gate applies. (src/agents/tool-search-runtime.ts:378, 3a6e3673c37a)
  • Regression coverage: The complete test file covers all three composition keywords, nested schemas, metadata refresh, visibility, tuple items, boolean/literal exclusion, bounded cycles, and both untrusted sources. Existing exact-name precedence remains covered; no new production test seam is added. (src/agents/tool-search-ranking.test.ts:225, 3a6e3673c37a)
  • After-fix production-path proof: The complete supplied PR body records before-fix misses and after-fix discovery through createToolSearchTools and actual catalog compaction in six mode/keyword combinations. Parameter-name and description queries find the tool; describe and valid calls succeed; invalid and revoked calls are denied. The committed-head run reports pass and an empty source diff. The embedded script was inspected, never executed by this reviewer. (3a6e3673c37a)

Likely related people:

  • steipete: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (1 earlier review cycle)
  • reviewed 2026-10-03T04:37:06.710Z sha 3a6e367 :: needs maintainer review before merge. :: none

@obviyus

obviyus commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Verified the actual built Gateway path on current main df883d4cc7bdbfc724cc167ccb2b2e2ecfd727cd and candidate 3a6e3673c37a520e6b03be24368ddf5b9c577f51, in isolated Linux containers with Tool Search enabled.

An explicitly allowed local plugin registers indexed_resource with payload.anyOf containing an object property named orchard, described as Collect apples. Neither word appears in the tool name or description. This is an OpenClaw plugin catalog entry, within the documented trusted-schema boundary.

Gateway tool_search query Current main Candidate
orchard [] indexed_resource
apples [] indexed_resource
saffron (flat-schema control) flat_resource, flat_secondary Same matches and order
xylophone (unused $defs behind recursive $ref) [], completes [], completes

The first root-level single-branch union fixture was normalized before cataloging and therefore already matched on main; the nested union above preserves the composition at the real indexing boundary and discriminates the fix.

The candidate regression file run against current-main production code had 7 failures and 34 passes (22.22s); the candidate passed all 41 tests (23.10s). This includes all three composition keywords, nested schemas, unchanged ranking cases, both untrusted sources remaining untraversed, and bounded cyclic traversal.

Live check on the candidate: two real gpt-5-mini requests through a passive proxy, both HTTP 200, reasoning_effort: low, max_completion_tokens: 4096. The model called tool_search for orchard, received indexed_resource, and reported the discovered tool without executing it. Utility jobs were off. An earlier proxy preflight rejected requests missing the explicit reasoning field before any provider call; configuring the model's supported reasoning capability resolved that fixture setup issue. No request/response payloads were rewritten.

No overlap with Pash/Sarah changes. No source changes were needed on top of @harshitgupta31415's fix, and the guarded push confirmed the existing head. No CI jobs were rerun. CI and ClawSweeper are ready on that exact SHA; the PR is mergeable.

@obviyus
obviyus merged commit 805a5be into openclaw:main Oct 3, 2026
245 of 250 checks passed
@obviyus

obviyus commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Landed as 805a5be. Thanks @harshitgupta31415, for the report and the fix!

github-actions Bot pushed a commit to Desicool/openclaw that referenced this pull request Oct 4, 2026
Fixes openclaw#164024.

Tool Search did not index a trusted tool's parameter names or descriptions when they sit under `anyOf`, `oneOf` or `allOf`, so queries matching only those terms found nothing. Composed schemas are now walked with bounded, cycle-safe traversal. Flat-schema matches and ranking order are unchanged, and untrusted MCP and client tools stay unindexed.

Proof: on a built Gateway, `tool_search` for terms inside a trusted plugin's nested `anyOf` returned nothing on base and `indexed_resource` on this change. A recursive `$ref` schema completed without indexing unused definitions. The regression tests fail on base (7) and pass here. In a live gpt-5-mini turn, the model called `tool_search` and found the tool.

Co-authored-by: Ayaan Zaidi <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agents Agent runtime and tooling P2 Normal backlog priority with limited blast radius. proof: sufficient ClawSweeper judged the real behavior proof convincing. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. size: S status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Tool Search misses parameters inside composed schemas

2 participants