Pular para o conteúdo
AvraAvra
Esc
↑↓navegar↵abrir⌘Jpré-visualizar
Nesta página

Migrating to Batch v2

The v2 Batch and Files API replaces single-file batch inputs with multi-file upload sessions and batches, improving speed, security and flexibility.

The core changes to the Batch API are related to how you get data in and results out: instead of uploading one file at a time you can now upload one or more files into a session, keeping subfolder organizations and reference them when submitting a batch job as avra://uploads/<session_id>/file.ext.

What changed

Concern v1 (deprecated) v2
Upload POST /v1/api/batch-inputs/upload — one file, referenced by UUID input_file_id POST /v2/api/files/upload — many files per session, integrity-checked with MD5
Discover files GET /v1/api/batch-inputs GET /v2/api/files with prefix / session / suffix filters
Create batch input_file_id + optional inputs, webhook_url input.files[] (avra:// paths), optional reference_date for backtesting
Result content_at URL on GET /v1/api/batches/{id}/result download_url + expires_at on GET /v2/api/batches/{id}/result

Webhooks are now configurable only through our notifications tool, via dashboard (see Webhooks). Input and Result deletions and Input downloads are no longer available in self-service ways to ensure security. Reach out to Avra if you need to remove a file.

Accepted file formats

  • Parquet — recommended for all inputs, and the format results are written in (result.parquet).
  • CSV — supported for low-complexity schemas.
  • JSONL — deprecated due to low compression. Migrate to Parquet.

Migration steps

Upload files into a session

Call POST /v2/api/files/upload with one target per file. Each target declares the destination file path, its content_type, and a base64-encoded md5 digest. You will receive one signed upload_url per target and a session expiration.

{
  "targets": [
    { "file": "base/input1.csv", "content_type": "text/csv", "md5": "KPRxZBtAn5UfexFEKfKr7Q==" }
  ]
}

For each target file, you will receive:

  • file: a reference for that file within Avra’s domain, avra://uploads/<session_id>/file.ext
  • content_type: a MIME type for that file, matching the Content-Type provided in request or, if not provided, inferred from the file extension.
  • md5: exact MD5 provided in the request as a counter-proof to use.
  • upload_url: a signed, short-lived URL to upload the described file, that expects Content-Type and Content-MD5 headers matching the response values. A mismatched digest will be rejected by the storage.
{
  "targets": [
      {
          "file": "avra://uploads/1784669990715/base/input1.csv",
          "content_type": "text/csv",
          "md5": "KPRxZBtAn5UfexFEKfKr7Q==",
          "upload_url": "https://storage.googleapis.com/..."
      }
  ],
  "expiration": "2026-07-21T22:39:50Z"
}

(Optional) List what you uploaded

Call GET /v2/api/files to enumerate uploaded files, optionally filtered by session, prefix, or suffix. Paths are returned in the avra:// scheme, ready to reference in a batch.

Create the batch

Call POST /v2/api/batches referencing the uploaded files. Provide model_id + model_version_id (or alias), an optional reference_date if backtesting, and the input definition.

{
  "model_id": "00000000-0000-0000-0000-000000000000",
  "model_version_id": "00000000-0000-0000-0000-000000000000",
  "reference_date": "2023-01-01",
  "input": { 
    "files": [
      "avra://uploads/1784143129686/base/input1.csv"
    ]
  }
}

Avra returns the batch id and status so you can track the job.

Poll and download

Poll GET /v2/api/batches/{id} for status. When complete, call GET /v2/api/batches/{id}/result to receive a signed download_url that expires at expires_at.

Esta página foi útil?