how to install @anthropic-ai/claude-agent-sdk and call query() for a stateless one-shot agent run
| Platform | Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 |
|---|---|
| Category | Automation Tools |
| Guide type | Procedure |
| Skill level | Beginner to intermediate |
| Time | 5 - 30 minutes including verification |
When how to install @anthropic-ai/claude-agent-sdk and call query() for a stateless one-shot agent run bites you on Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026, the first instinct is to rerun the whole scenario or redeploy the script. Most of the time you do not have to. The steps below are what an automation engineer would do at their desk before escalating - The pattern I see most often is in Make so the working state is always reproducible by branch.
What how to install @anthropic-ai/claude-agent-sdk and call query() for a stateless one-shot agent run actually involves on Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026
On Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 in my experience the most useful first-pass tools are Honeycomb or Jaeger UI for span inspection, pytest -k agent_sdk with VCR.py for replay fixtures, npm ls @anthropic-ai/claude-agent-sdk to confirm pinned version. Each of these surfaces a different layer of the failure - keep at least the first one in your personal notes so the next time this happens you do not start cold.
For verification on Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026, the methods that survive contact with a real Monday-morning workload are otel-cli exporter probe to confirm trace ingestion and pip install claude-agent-sdk && python -c "import claude_agent_sdk; print(claude_agent_sdk.__version__)". Anything less than that and you are shipping on vibes.
Authoritative sources for Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 that I cross-reference before committing to a fix: docs.anthropic.com, github.com/anthropics/claude-agent-sdk-python, github.com/anthropics/skills. Marketing blog posts and Medium writeups are signal, not ground truth.
The rest of this page is the structured fix path. Start with diagnose, then remediation, then the automation options so you do not have to do this by hand the next time it surfaces. Verify and safety sections at the end are the discipline that keeps the fix from regressing the next time you open the platform.
Diagnose first, fix second
Sixth: pin down the latency and reliability envelope on the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 session under real working conditions. Run a long-duration sanity test by executing the failing scenario 10 times over 15 minutes, logging the timestamp and the result (success / error code / which step failed) per attempt to a notes file. Watch for the breakpoint where the success rate dips below 80 percent - that is your real signal that something is wrong, not the one-off failure that prompted the investigation. If you are on a marginal network (cafe wifi, mobile hotspot, hotel network), run the same test on a wired or known-good connection before assuming the platform is the problem. Capture the breakpoint in your personal notes next to the platform version, the account, and the workspace id - the next time this happens to a teammate, the notes are gold.
Fifth: replay the failing run against a second account or a second connector on the same Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 workspace. The point is to isolate "my credentials" from "my account" from "the whole workspace." If a teammate's identical scenario works but yours does not, the failure is local cache or a stale OAuth grant. If the same scenario fails for everyone in the same workspace, you have a tenant-wide config change or a vendor-side incident. Pin the platform version explicitly while you do this: the platform's About panel, the build hash in the footer, or the engine version returned by a diagnostic call. The version pin is what isolates "their rollout broke me" from "my client is out of date."
Fourth: open the vendor status page for Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 and the connector's upstream status pages for the failing window. The smoking guns are an open incident touching the exact service area you are using, a recent post-mortem covering the same symptom, or a Trust Center advisory on a partial outage. Cross-reference the timestamp of your first failed run against the incident start time - if they match within 5 minutes, stop debugging your own setup and subscribe to the incident updates. Many vendors lag the status page behind the actual incident by 10 to 30 minutes; if Twitter and Reddit are both lit up but the status page is green, trust the crowd and treat it as upstream until proven otherwise.
Field notes from real Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 incidents
On any Agentic AI problem in Claude Agent SDK, the first three questions I ask are: which runtime, which tenant, which trigger source. Defaults shift quietly between platform updates. My go-to verification step is `claude /agents and confirm the skill pack folder is recognized when SDK and CLI share a project`; I learned the hard way that the Claude Agent SDK UI will happily lie about whether a flow really ran.
In Agentic AI work, the cost of guessing is almost always higher than the cost of reading the Claude Agent SDK changelog, read the changelog first. I keep Honeycomb or Jaeger UI for span inspection docked on a second screen whenever I am building inside Claude Agent SDK; one glance tells me whether the run actually fired or silently skipped.
Tools I actually reach for
For most Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 stalls I start with uv pip show claude-agent-sdk for Python install metadata, fall back to claude-agent-sdk-python verbose logger via logging.DEBUG, Anthropic Workbench for prompt regression baselines when uv pip show claude-agent-sdk for Python install metadata cannot surface the answer, and keep Anthropic Console message logs panel handy for the cases where neither answers. That ordering is not academic - it matches the layers of the failure as they tend to surface, so the cheapest signal lands first and the heavier tooling only comes out when the simpler answer does not hold up. My muscle-memory shortcut for this is to run the first tool while the failing screen is still open, not after I have already restarted the platform.
Verification I run before I call it fixed
Before I mark a Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 stall resolved, the verification loop below is what I actually run. Each step proves a different layer is green, and the order matters - the cheaper checks gate the more expensive ones.
claude /agents and confirm the skill pack folder is recognized when SDK and CLI share a projectIf that one comes back clean, move to the next check. If it does not, stop and dig in there before layering more verification on top of a red signal.
npm install @anthropic-ai/claude-agent-sdk && node -e "import('@anthropic-ai/claude-agent-sdk').then(m=>console.log(Object.keys(m)))"If that one comes back clean, move to the next check. If it does not, stop and dig in there before layering more verification on top of a red signal.
pytest tests/agent_sdk --record-mode=none to enforce VCR replayOnly when every line above runs clean do I close the loop and update my notes with the timestamps.
Where I check first when the docs disagree
When two sources contradict each other on a Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 detail, the disambiguation order I lean on is stable. I usually check github.com/anthropics/skills for the ground-truth view on this part of Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026. I usually check code.claude.com/docs/en/agent-sdk/overview for the ground-truth view on this part of Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026. I usually check docs.anthropic.com for the ground-truth view on this part of Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026. Marketing blog posts and Medium writeups are signal, not ground truth, and I treat them as such until the references above either confirm or contradict the claim.
Solution-focused remediation path
Before any destructive step on a Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 workspace, slow down and stage rollback. Snapshot the current platform version, the current workspace settings (Settings -> screenshot every tab), the connected-apps list, the current sharing policy, and the current member list to a notes entry first. Capture the failing screenshot, the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 incident id if any, and the timestamp window. Photograph (screenshot) the workspace state from two angles: the scenario or script that is failing, and the workspace settings page that controls the relevant policy. Then do the destructive step (revoke a connector, change a sharing default, remove a member, delete a connected app) inside a test workspace or a test scenario first, never the whole workspace. Capture the platform version, the API permissions, the connected-app list, the workspace member roster, and the relevant integration log snapshot to your notes before the destructive step. Decision point: if you are on a paid plan, the cheapest correct path is almost always to open the in-product support chat in parallel with the rollback - the support rep can confirm whether a vendor-side rollout is responsible while you are still staging the change, which avoids a needless workspace edit if the fix is server-side.
For Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 integrations where rate limits or plan quotas are suspect, read the in-product hints honestly. "You have reached the limit for this workspace" usually means you hit an operation, task, or run cap on the current plan tier. "Slow down, you are sending requests too quickly" is the rate-limit signal on the trigger source or destination API. "This payload is too large" is the per-call cap. Each is telling you the exact same thing in a Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026-specific dialect. Apply exponential backoff for API-driven runs (base 1s, double up to 60s, retry up to 5 times) and split a large batch into chunks of 100 records at a time. Decision point: if you are hitting the quota sustained rather than in bursts, upgrade the plan tier or request a quota increase from the workspace admin with a written usage justification; without it, batch the work or shed load at the producer. Replay the failing scenario against a fresh test workspace at half the throughput to confirm the new safe rate before pushing to the real workspace.
When the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 fault tracks to integration failures, automation delays, or webhook drops from the trigger source (the trigger source, the connector, the upstream provider), treat the integration plane as suspect. Open the integration log in the connected service (the trigger source's webhook log, the platform's connector run history) and read the response status the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 endpoint actually returned - most "scenario not firing" reports are actually "webhook firing but the connector failed and the platform backed off." Verify the connected account is still authorized (the OAuth grant in Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 is not silently revoked) and that the trigger event is what you think it is. Decision point: if the trigger is firing but Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 is rate-limiting it, throttle the scenario (bump the polling interval, add a sleep module, enable batch mode) and re-run. Verify the connected workspace is the right workspace - a common foot-gun is the personal workspace being authorized while the work workspace holds the data.
Automate this fix so you do not do it twice
Codify the platform version pin and rollback as a single notes entry
Once a stable platform version is identified for the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026, write the version string, the build hash, and the workspace policy state to a personal notes entry with the date in the title. Reproducible rollback is then a single download-and-install plus a sign-in. Pin the workspace policy state explicitly so a vendor-side default change does not silently shift behavior under you. Stage the notes entry next to a checklist that lists the failing screenshot, the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 incident id (if any), and the support case number; the second time the workflow breaks at 9 a.m. you do not want to be rediscovering which platform build was actually green.
# Personal notes template (claude)
Date: 2026-05-31
Platform: claude
Working build: 2.45.1 (Build hash: a1b2c3d)
Account: [email protected]
Workspace: ws-prod-claude
Failing screenshot: ~/notes/claude-2026-05-31.png
Support case: SUPP-claude-12345
Rollback path: download installer from vendor releases page, sign out, reinstall, sign back inAutomate Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 session + sharing-policy snapshots via vendor CLI or API
On the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026, regular session and policy snapshots catch silent role changes, sharing-default drift, and stale OAuth grants well before the workflow starts failing in prod. Pair vendor health checks (the platform's admin SDK, the platform's users API, the connector listing) with a token-validity check so both vendor-side and account-side issues land in one folder. Run the scheduled task on a control plane device (a small VPS, a GitHub Actions runner, a Cloud Function) under a tightly scoped service account that mirrors the real workspace policy.
# List workspace members + roles
curl -H "Authorization: Bearer $PLATFORM_TOKEN" \ https://api.example.com/v1/workspace/members \ > claude-members.json
# List active connectors + their last-tested timestamp
curl -H "Authorization: Bearer $PLATFORM_TOKEN" \ https://api.example.com/v1/connectors \ > claude-connectors.json
# Validate the bearer token itself
curl -H "Authorization: Bearer $PLATFORM_TOKEN" \ https://api.example.com/v1/me \ > claude-me.jsonMonitor + alert via Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 admin reports, audit logs, and personal dashboard ingestion
For the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026, the most useful long-running telemetry is the admin reports + audit logs shipped to a personal dashboard (Google Sheets daily import, Airtable scheduled sync, Notion database via the API, Grafana with a CSV source) and graphed on a single view. Pair that with synthetic monitoring (a small script that triggers the failing scenario or runs the failing action every 5 minutes from at least two devices) so a regional incident lights up before teammates report it. Subscribe the personal inbox or a private Slack channel to the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 status page (Atom/RSS or Statuspage webhook) plus the vendor X/Twitter status handle so an open incident self-correlates with the synthetic failures.
# Tiny synthetic monitor - hit the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 health endpoint every 5 minutes
while true; do curl -s -o /dev/null -w "%{http_code} %{time_total} $(date -Iseconds)\n" \ -H "Authorization: Bearer $TOKEN" \ https://api.example.com/v1/me \ >> ~/logs/claude-synth.log sleep 300
done
Common pitfalls and what to watch for
Read-only validation before any write is the single step most Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 fixes skip, and it is the step that lets you roll back when a fix backfires. Screenshot every existing settings page (the workspace settings, the sharing policy, the connected-apps list, the members page, the plan tier page), capture the failing screenshot in a notes entry, export the relevant log to CSV if the platform supports it (the platform's run-history export, the audit-log download), and screenshot the activity feed showing the failing window before any change. On Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 workspaces with multiple environments (test workspace, real workspace) record the platform version, the settings state, and the connected-apps list in each before toggling anything, because a "fix" pushed only to the test workspace is a known regression vector when the real workspace has a different policy.
The mirror-image mistake is confusing a user-side symptom with a vendor fault on Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026. A persistent 403 is often a connector-level change pushed by the workspace owner rather than a Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 bug. A "scenario not found" can be a moved scenario rather than a deleted one. A "webhook not firing" is frequently a corporate proxy or firewall dropping the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 egress IP rather than a vendor-side regression.
Verify the fix worked
- Reproduce the original failing run against Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 on the same device AND a second device with the same account. If the failing toast or error code still surfaces on any device, you have not fixed it.
- Watch for 24 to 48 hours via the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 workspace audit log + the integration history + your personal notes. Cached error states and CDN caches mask slow-burn drift and intermittent regional issues.
- Smoke-test under realistic load: replay the workflow against a test workspace for at least 30 minutes at your normal working pace, log success / error and the timestamp per attempt to a notes file.
- Capture the new state in a personal notes entry so the next time this happens you do not rediscover it. Note platform version + workspace policy + connected-apps list + failing screenshot + verbatim error string + fix applied. Push to a shared team wiki if your team uses one.
- If the fix involved an API token rotation or a workspace policy change, commit the new token to your password manager and screenshot the workspace settings for archival.
Safety, rollback, blast radius
- Test in a Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 test workspace or on a duplicate scenario first before any change that touches the real workspace. Snapshot the platform version, the workspace settings, the connected-apps list, and the sharing policy before changing anything.
- Apply the principle of least surprise when granting share access or connected-app permissions. Review the share list against the people who actually need access - extra shares are extra blast radius.
- Use idempotent runs where the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 API supports it (the platform's run id de-dupe, external id keys on destination records) so a retried run does not create duplicate records.
- Know your rollback path. Platform version rollback is a one-line download-and-install; an API token rotation is reversible if you kept the old token in the password manager during cutover; a workspace policy change is reversible only if you saved the previous policy in a screenshot.
- For team-wide or workspace-wide changes, line up a maintenance window with team notification before pushing through the admin console.
FAQ
References
- Vendor help center for Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 (official help articles, API docs, Trust Center)
- Community forums (r/nocode, r/automation, r/GoogleAppsScript, r/PowerAutomate, r/n8n, r/make, r/ClaudeAI, vendor community)
- In-product help and the Claude Agent SDK - Skill Packs, Tool Use, Evaluation Harnesses - 2026 changelog
- Vendor status pages and X/Twitter status handles, plus post-mortem incident reports
Related fixes
Related guides worth a look while you sort this one out:
- how to add a pre-tool-use hook in the Agent SDK that vetoes a Write call to a protected path
- how to export OpenTelemetry traces from the Agent SDK to Honeycomb or Jaeger for tool-call inspection
- how to fix Agent SDK 401 from the Anthropic API when ANTHROPIC_API_KEY is set but unread by the subprocess
- how to fix Agent SDK query() returning empty assistant messages because permission_mode rejected the tool
- how to measure tool-call success rate as an Agent SDK harness metric across a benchmark set
- how to build an evaluation harness that scores Agent SDK runs against a golden-trajectory dataset