Best Practices
Ask sharper questions, pivot efficiently across datasets, and validate what your assistant finds.
Early AccessSpyCloud Investigations MCP is in Early Access. Tools, limits, and behavior may change before general availability, targeted for early 2027.
Your assistant decides which SpyCloud tools to call, how to combine them, and how to present the results. The more specific your request, the faster, cheaper, and more reliable the investigation. Every tool call draws one query from your Investigations API key.
Ask specific questions
| Instead of | Try |
|---|---|
| "Find exposures." | "Find breach records for [email protected]." |
| "Is this domain bad?" | "Show workforce exposure for example.com over the last 24 weeks, with the password reuse rate." |
| "Look into this machine." | "List every infostealer log for infected machine ID <id>, then show collection sizes for the newest log." |
For better results:
- Name the asset. Email, domain, username, phone, IP address, log ID, or infected machine ID.
- Add a time range when recency matters.
- Say what you want back. "Return the raw records rather than a summary," or "Show the results as a table."
- Ask for the tool calls when you need to repeat or audit the work: "Show which SpyCloud tools you used and the parameters for each call."
- Narrow before you widen. If a query is slow or returns too much, tighten the asset, dataset, or date window.
Follow the investigation
Identifiers returned by one tool open the door to others. These pivots cover most cases.
| Start with | Call | To learn |
|---|---|---|
source_id on a breach record | breach_catalog_get | Which breach or source the record came from |
document_id on a breach record | breach_data_doc_ids | The full record |
infected_machine_id | infostealer_logs_by_machine | Every log from that device, including repeat infections and multiple malware families |
log_id | infostealer_log_metadata, then infostealer_log_collections | What the infection captured |
| Email, phone, username, or handle | idlink_query at depth 1, then idlink_nodes filtered by minimum confidence | Connected identities, accounts, and devices |
log_id or infected_machine_id from breach_data | idlink_query | Other identities tied to the same infected device |
| Domain | domain_stats, then domain_timeseries, then botnet_customers | Workforce exposure, its trend, and customer-side infostealer impact |
Work an infostealer log in order
Call the infostealer tools in this sequence:
infostealer_logs_by_machineshows every log for the device, so you can spot repeat infections before you commit to one log.infostealer_log_metadatashows what the chosen log holds and how large each collection is.infostealer_log_collectionspulls only the collections you need.
Sizing a log before you pull it keeps large collections from overwhelming your assistant's context.
Expand the identity graph gradually
Start idlink_query at depth 1. Review the top pivots and node counts, then use idlink_nodes with a minimum confidence to retrieve only the strongest connections. Move to depth 2 only when depth 1 leaves a clear gap; a depth 2 graph can return far more data.
Keep large investigations manageable
- Start broad, retrieve selectively. Use
breach_statsordomain_statsto size an exposure before you pull records. - Use summary detail first. Set
detail_leveltosummaryonbreach_data, then pull full records only for the hits that matter. - Handle big collections outside the conversation. For collections above 1,000 items, such as large cookie stores, have your assistant write results to a file if your client supports it, then filter with
jq,grep, or Python. Bring the relevant subset back into the conversation. - Cap cookies. Use
cookie_limitoninfostealer_log_collectionswhen a log has thousands of cookies.
Read the data correctly
- Publish date is not infection date.
spycloud_publish_daterecords when SpyCloud acquired the data. For redistributed logs, that can be years after the malware ran. Checkinfected_timebefore describing activity as current. - A machine ID lookup can miss logs. If a log's
infected_machine_idis empty,infostealer_logs_by_machinecannot find it. Recover thelog_idthroughbreach_data, then query the log directly. - One log, many collections. Credentials and cookies often arrive in the same log. Check
infostealer_log_metadatabefore treating them as separate sources. - Cookies are the most valuable and most often skipped. Some session tokens survive a password reset. If you remediate by resetting passwords alone, those sessions stay live. Revoke sessions as well.
- Autofills carry more than credentials. They can include addresses, payment details, and anything the user typed into web forms.
- Domain tools measure different populations. Workforce exposure (
domain_stats) and customer exposure (botnet_customers) will not match, and they shouldn't. - No results is not proof of no exposure. It means SpyCloud found no matching records for that selector.
- Match the selector format. Username search is case-sensitive. Remove the leading
+from phone numbers.
Validate before you act
Your assistant's answer is a synthesis of what SpyCloud's tools returned. When the evidence matters, look behind the summary:
- "Which SpyCloud records support that finding?"
- "Show me the source metadata for that breach."
- "Show the underlying records, not just the summary."
MCP results are AI-assisted and are not a system of record. Before a finding goes into a case file, report, or legal filing, confirm it with a deterministic pull from the Investigations API.
What to expect
- Response times vary. Large breach searches can take several seconds. Identity graph and statistics calls are usually faster.
- Large results may be summarized. Ask for the underlying data explicitly when you need it.
- Clients render results differently. Cursor generally renders tables more fully. Claude Code and Claude Desktop are more text-forward.
- Data is current. MCP tools query the same production data as the Investigations API. New records become available to the MCP as they're published.
- Similar questions can take different paths. Your assistant composes tool calls dynamically, so the same question may produce a different sequence. For repeatable results, review the tool calls or use the Investigations API.
Updated about 2 hours ago