Best Practices

Ask sharper questions, pivot efficiently across datasets, and validate what your assistant finds.

🚧

Early Access

SpyCloud Investigations MCP is in Early Access. Tools, limits, and behavior may change before general availability, targeted for early 2027.

Your assistant decides which SpyCloud tools to call, how to combine them, and how to present the results. The more specific your request, the faster, cheaper, and more reliable the investigation. Every tool call draws one query from your Investigations API key.

Ask specific questions

Instead ofTry
"Find exposures.""Find breach records for [email protected]."
"Is this domain bad?""Show workforce exposure for example.com over the last 24 weeks, with the password reuse rate."
"Look into this machine.""List every infostealer log for infected machine ID <id>, then show collection sizes for the newest log."

For better results:

  • Name the asset. Email, domain, username, phone, IP address, log ID, or infected machine ID.
  • Add a time range when recency matters.
  • Say what you want back. "Return the raw records rather than a summary," or "Show the results as a table."
  • Ask for the tool calls when you need to repeat or audit the work: "Show which SpyCloud tools you used and the parameters for each call."
  • Narrow before you widen. If a query is slow or returns too much, tighten the asset, dataset, or date window.

Follow the investigation

Identifiers returned by one tool open the door to others. These pivots cover most cases.

Start withCallTo learn
source_id on a breach recordbreach_catalog_getWhich breach or source the record came from
document_id on a breach recordbreach_data_doc_idsThe full record
infected_machine_idinfostealer_logs_by_machineEvery log from that device, including repeat infections and multiple malware families
log_idinfostealer_log_metadata, then infostealer_log_collectionsWhat the infection captured
Email, phone, username, or handleidlink_query at depth 1, then idlink_nodes filtered by minimum confidenceConnected identities, accounts, and devices
log_id or infected_machine_id from breach_dataidlink_queryOther identities tied to the same infected device
Domaindomain_stats, then domain_timeseries, then botnet_customersWorkforce exposure, its trend, and customer-side infostealer impact

Work an infostealer log in order

Call the infostealer tools in this sequence:

  1. infostealer_logs_by_machine shows every log for the device, so you can spot repeat infections before you commit to one log.
  2. infostealer_log_metadata shows what the chosen log holds and how large each collection is.
  3. infostealer_log_collections pulls only the collections you need.

Sizing a log before you pull it keeps large collections from overwhelming your assistant's context.

Expand the identity graph gradually

Start idlink_query at depth 1. Review the top pivots and node counts, then use idlink_nodes with a minimum confidence to retrieve only the strongest connections. Move to depth 2 only when depth 1 leaves a clear gap; a depth 2 graph can return far more data.

Keep large investigations manageable

  • Start broad, retrieve selectively. Use breach_stats or domain_stats to size an exposure before you pull records.
  • Use summary detail first. Set detail_level to summary on breach_data, then pull full records only for the hits that matter.
  • Handle big collections outside the conversation. For collections above 1,000 items, such as large cookie stores, have your assistant write results to a file if your client supports it, then filter with jq, grep, or Python. Bring the relevant subset back into the conversation.
  • Cap cookies. Use cookie_limit on infostealer_log_collections when a log has thousands of cookies.

Read the data correctly

  • Publish date is not infection date. spycloud_publish_date records when SpyCloud acquired the data. For redistributed logs, that can be years after the malware ran. Check infected_time before describing activity as current.
  • A machine ID lookup can miss logs. If a log's infected_machine_id is empty, infostealer_logs_by_machine cannot find it. Recover the log_id through breach_data, then query the log directly.
  • One log, many collections. Credentials and cookies often arrive in the same log. Check infostealer_log_metadata before treating them as separate sources.
  • Cookies are the most valuable and most often skipped. Some session tokens survive a password reset. If you remediate by resetting passwords alone, those sessions stay live. Revoke sessions as well.
  • Autofills carry more than credentials. They can include addresses, payment details, and anything the user typed into web forms.
  • Domain tools measure different populations. Workforce exposure (domain_stats) and customer exposure (botnet_customers) will not match, and they shouldn't.
  • No results is not proof of no exposure. It means SpyCloud found no matching records for that selector.
  • Match the selector format. Username search is case-sensitive. Remove the leading + from phone numbers.

Validate before you act

Your assistant's answer is a synthesis of what SpyCloud's tools returned. When the evidence matters, look behind the summary:

  • "Which SpyCloud records support that finding?"
  • "Show me the source metadata for that breach."
  • "Show the underlying records, not just the summary."

MCP results are AI-assisted and are not a system of record. Before a finding goes into a case file, report, or legal filing, confirm it with a deterministic pull from the Investigations API.

What to expect

  • Response times vary. Large breach searches can take several seconds. Identity graph and statistics calls are usually faster.
  • Large results may be summarized. Ask for the underlying data explicitly when you need it.
  • Clients render results differently. Cursor generally renders tables more fully. Claude Code and Claude Desktop are more text-forward.
  • Data is current. MCP tools query the same production data as the Investigations API. New records become available to the MCP as they're published.
  • Similar questions can take different paths. Your assistant composes tool calls dynamically, so the same question may produce a different sequence. For repeatable results, review the tool calls or use the Investigations API.

Did this page help you?