Skip to main content

File and URL analysis providers

AiSOC talks to malware sandboxes and file-reputation services through one interface, so the console, the enrichment path and the investigation agent all read the same words whichever backend answered. A verdict, a score, signatures, IOCs and an ATT&CK mapping are the vocabulary; each provider maps its own payload onto them.

Three providers ship today:

ProviderRunsUploadsStatus
capev2Your own networkYesOpen-source reference. Written against the documented CAPEv2 REST API, unverified against a live instance
malwareanalyzerHosted, third partyYes, published by defaultCommercial adapter. Hash lookup and report mapping verified against the live service; authentication unverified
mockIn processNo effectAnalyses nothing. Present so the wiring is exercised on a deployment with no sandbox configured

The rule that matters: uploading a file is a disclosure​

A SHA-256 lookup tells a provider a 32-byte digest. Uploading the file tells it the file. Those are different acts, and AiSOC treats them differently.

Hash lookup always runs first. If the provider has already analysed the file, the existing report is used and nothing is uploaded. This is the common case and it costs nothing.

Uploading is off by default. It must be enabled per tenant and per provider, by someone holding settings:write, and the decision is recorded with who made it and which version of the disclosure text they were shown. Consent is per provider because consenting to a sandbox in your own rack says nothing about a hosted service.

Air-gapped mode permits local providers only. When AISOC_AIRGAPPED is on, a hosted provider stays visible in the settings list, marked excluded with the reason, rather than silently disappearing.

Every decision is logged, refusals included. aisoc_sandbox_submissions holds one row per artefact with the outcome, the reason and an uploaded boolean that is true only when bytes actually left the deployment. That column is what answers "did any customer file go to a third party".

An operator enabling uploads to malwareanalyzer is shown this:

Enabling uploads lets AiSOC send files from this tenant to malwareanalyzer for analysis. A file may contain customer data, credentials, personal data or intellectual property. AiSOC cannot inspect a file to decide whether it does. malwareanalyzer publishes submitted samples by default: an uploaded file and its report become readable by anyone on the internet, not only by the vendor. Treat every upload as a public disclosure that cannot be recalled. AiSOC always looks a file's SHA-256 up first and only uploads when the provider has not seen it. Every upload and every refusal is recorded against this tenant.

The word public is there because it is literally what the service reports. A stored report fetched from GET /v1/reports/<sha256> carries visibility: "public" and tlp: "clear". TLP:CLEAR is the Traffic Light Protocol's unrestricted level, so this is not a vendor-internal default: the sample is readable by anyone.

The service's own web client does send a visibility field on submission and offers a private option, but it pins public and disables the control when there is no signed-in account. So private appears to need an authenticated account. AiSOC requests the visibility you configure and assumes nothing about whether it is honoured: the consent text keeps saying public until an operator sets MALWAREANALYZER_PRIVATE_SUBMISSIONS_CONFIRMED=1, which is a statement that they confirmed private submissions work for their account. Letting a mere API key flip that wording would be an unverified assumption rewriting the sentence somebody agrees to before disclosing a customer file.

Configuration​

Nothing is enabled by default. A hosted provider that switched itself on because a base URL had a default would be a deployment making outbound calls nobody asked for.

AISOC_CAPEV2_URL=http://cape.internal:8000
AISOC_CAPEV2_TOKEN= # optional; omit entirely on an open instance

Leave the token unset if your instance does not require one. AiSOC sends no Authorization header at all rather than an empty one, which would turn an open instance into a 401.

MalwareAnalyzer​

AISOC_MALWAREANALYZER_ENABLED=1
MALWAREANALYZER_BASE_URL=https://malwareanalyzer.com
MALWAREANALYZER_API_KEY= # optional, see below
MALWAREANALYZER_API_KEY_HEADER=Authorization
MALWAREANALYZER_API_KEY_PREFIX="Bearer "
MALWAREANALYZER_PRIVATE_SUBMISSIONS_CONFIRMED= # only after you have verified it
Authentication is unverified

The service publishes "No API key required", and every call AiSOC makes works without one. There is an authenticated surface, and the service's own web client sends Authorization: Bearer <token>, which is why that is the default here. But a bogus Bearer and a bogus x-api-key produce byte-identical 401s, and the CORS preflight reflects whatever header is asked for, so neither confirms the scheme for a minted API key.

Nobody has confirmed this against a real key. MALWAREANALYZER_API_KEY_HEADER and MALWAREANALYZER_API_KEY_PREFIX exist so you can correct it without a code change, and the settings surface reports authentication as unverified rather than showing a tick this project has not earned. If you hold a real key and determine the scheme, please open an issue.

How a pending or failed analysis surfaces​

A sandbox is slow and failure-prone. Two states are kept distinct from a clean result everywhere they travel:

  • Pending. A submitted file has no verdict yet. The report carries state: pending or running and its verdict field reads unavailable. It is never rendered as benign, and the agent is told the analysis is incomplete.
  • Could not check. A provider that timed out, refused authentication or returned something unparseable produces outcome: could_not_check with the reason attached. It never produces a report.

To the investigation agent, both arrive as available: false with wording that says plainly that the file was not assessed and that this is not evidence the file is benign. That distinction is the one that matters: a model reading a timeout as "no detections" writes a benign verdict on a file nobody analysed.

An unknown hash is handled the same way. "No provider has ever analysed this file" is genuinely different from "a provider analysed it and found nothing", and targeted malware is unknown to every public service by design.

"Unavailable" is not zero​

Where a provider does not publish one of the interface's fields, it reads unavailable, never 0 and never [], and the report says why.

The clearest example is in the test fixtures. A MalwareAnalyzer report for a file whose type could not be identified carries attackTechniques: [] next to behavior.analyzed: false. The list is empty because no guest was ever chosen and the behavioural stage never ran, not because the sample exhibits no techniques. AiSOC reports that as:

{
"attack": "unavailable",
"unavailable_reasons": {
"attack": "behavioural analysis did not run (unidentified_file_type), so no techniques could be observed"
}
}

The same payload with the behavioural stage marked as having run produces a real empty list, which means "we looked and found none".

The agent tool​

The investigation agent gets one tool, lookup_file_hash. It cannot upload.

That is deliberate, not a transport limitation. An upload is a disclosure governed by a consent an operator recorded deliberately, and a decision that consequential does not belong behind a sentence a model chose to emit. An injected instruction in an attachment name would otherwise be one step away from exfiltrating the attachment.

The tool authenticates with the agent service's own API key (AISOC_AGENTS_API_KEY) and passes no tenant. The API resolves the tenant from that credential, so a prompt cannot redirect the lookup or reach another tenant's history.

API​

RoutePermissionNotes
GET /api/v1/sandbox/providersthreat_intel:readProviders, their consent state and the current disclosure text
POST /api/v1/sandbox/lookupthreat_intel:readHash lookup. Never uploads
POST /api/v1/sandbox/filesthreat_intel:writeHash lookup, then upload only if consent exists and confirm_upload is set
POST /api/v1/sandbox/urlsthreat_intel:writeURL submission. Not covered by the upload policy, still covered by air-gap
GET /api/v1/sandbox/analyses/{provider}/{handle}threat_intel:readPoll
PUT /api/v1/sandbox/consentsettings:writeRecord or withdraw consent

POST /api/v1/phishing/submit accepts attachment_hashes, which are looked up through the same path. It accepts digests only: that route runs unattended on submitted mail, and an unattended path that can upload is one misconfiguration away from disclosing every attachment a tenant receives.

Adding a provider​

Implement SandboxProvider in services/api/app/services/sandbox/providers/, declare ProviderCapabilities including local and submissions_are_public_by_default, and register it in registry.py. Nothing outside that package should learn the new provider's name.

scripts/check_sandbox_upload_policy.py fails if an adapter does not declare its locality, if a consent-shaped flag defaults to on, if the hash lookup stops preceding the upload, or if an exception handler starts producing a report.