Current situation
IBM Storage Deep Archive computes and stores an MD5 for each object at ingest. It recomputes it and compares it with the stored value when the object is restored from tape. This value cannot be read through the S3 API, and it is not documented in the 1.3.1 documentation. The ETag returned by HeadObject comes from file-system attributes (e.g. "mtime-…-ino-…"), not from the object content, even for single-part uploads.
Today there are only two ways to get the MD5, and neither is practical for automated workflows:
- a storage administrator connects to a node over SSH and runs an internal command;
the application (MAM) runs RestoreObject and GetObject, then recomputes the MD5 on the client.
The second option means recalling the full object from tape and moving it over the network just to check a checksum. On a tape archive that costs a lot of time, drive usage and bandwidth.
IBM has said it plans to expose the MD5 in a dedicated HeadObject response header in the next release, similar to the existing x-tape-meta-copy-N headers. We welcome this and ask that it be delivered with the requirements below.
Request
- Expose the MD5 through the S3 API (next release)
- Return the MD5 of the complete object in a documented HeadObject (and GetObject) response header, independent of the ETag.
The header must be available for archived objects without a prior restore, as the tape barcode headers already are.
It must cover both single-part and multipart uploads, and give the MD5 of the full multipart object, not the composite multipart ETag.
Document the header name, its encoding (hex or base64) and its behaviour.
Validate Content-MD5 on upload
- Validate Content-MD5 on PutObject and UploadPart, and reject mismatches with
400 BadDigest.
Document this behaviour.
Make restore-time failures visible to S3 clients
- When the MD5 recomputed during restore does not match the stored value, the S3 client should see it: a clear error on RestoreObject/GetObject or in the Restore status, plus a system event or alert.
The object must never be served silently.
Support the standard S3 checksums (x-amz-checksum-*)
- Algorithms: SHA-256, SHA-1 and CRC32/CRC32C, with CRC64NVME optional.
Validate on upload (x-amz-checksum-*, x-amz-sdk-checksum-algorithm, trailing checksums) and return BadDigest on mismatch.
For multipart uploads, support per-part checksums and the checksum of the complete object (FULL_OBJECT and/or COMPOSITE).
Return the checksums through HeadObject/GetObject with x-amz-checksum-mode: ENABLED, through GetObjectAttributes and through ListParts.
Persist them with the object across tape migration and multi-copy/multi-site replication, and re-verify them on restore, as MD5 is today.
Rationale
Multipart upload will be the main ingest method for large media files. Applications such as Telestream DIVA already support SHA-1, SHA-256 and CRC32. If Deep Archive exposes the internal MD5 now, and adds standard S3 checksums later, applications can check integrity end to end without recalling objects from tape. They can also use the same checksum mechanism on every S3 platform they work with.