Skip to Main Content
IBM System Storage Ideas Portal


This portal is to open public enhancement requests against IBM System Storage products. To view all of your ideas submitted to IBM, create and manage groups of Ideas, or create an idea explicitly set to be either visible by all (public) or visible only to you and IBM (private), use the IBM Unified Ideas Portal (https://ideas.ibm.com).


Shape the future of IBM!

We invite you to shape the future of IBM, including product roadmaps, by submitting ideas that matter to you the most. Here's how it works:

Search existing ideas

Start by searching and reviewing ideas and requests to enhance a product or service. Take a look at ideas others have posted, and add a comment, vote, or subscribe to updates on them if they matter to you. If you can't find what you are looking for,

Post your ideas
  1. Post an idea.

  2. Get feedback from the IBM team and other customers to refine your idea.

  3. Follow the idea through the IBM Ideas process.


Specific links you will want to bookmark for future use

Welcome to the IBM Ideas Portal (https://www.ibm.com/ideas) - Use this site to find out additional information and details about the IBM Ideas process and statuses.

IBM Unified Ideas Portal (https://ideas.ibm.com) - Use this site to view all of your ideas, create new ideas for any IBM product, or search for ideas across all of IBM.

ideasibm@us.ibm.com - Use this email to suggest enhancements to the Ideas process or request help from IBM for submitting your Ideas.

Status Submitted
Created by Guest
Created on Oct 8, 2026

Reference architecture for resilient S3 access and failover across multiple IBM Storage Deep Archive systems

IBM Storage Deep Archive can be deployed with multiple interconnected tape libraries/systems that synchronize data between sites. In such an architecture, an S3 client may potentially connect to more than one system, providing opportunities for geographical traffic distribution, redundancy, resilience, and disaster recovery.

However, the current IBM Storage Deep Archive documentation does not appear to describe a reference architecture for providing resilient S3 client access across multiple replicated systems, nor does it clearly define the recommended mechanism for automatically failing over S3 traffic when the preferred system becomes unavailable.

For the purposes of this request, a “node” is considered a single TapeCloud server belonging to a tape library; a “system” or “tape library” is a unit consisting of two TapeCloud nodes and the associated robotics; and an “array” is a resilient configuration consisting of two or more interconnected tape libraries. In such an array, data is synchronized between the interconnected systems, allowing an S3 client to potentially access any of them.

This architecture raises an important operational question: how should the S3 endpoint presented to clients be selected, and how should the client automatically fail over when the preferred system is unavailable?

There are several possible architectural approaches.

Client-side endpoint selection

The S3 client could be configured with a preferred endpoint and one or more secondary endpoints. The client would be responsible for determining endpoint availability and switching to an alternative system when the preferred endpoint cannot serve requests.

This approach has the advantage of keeping the S3 data path direct, without introducing additional infrastructure. It is particularly attractive when the S3 client already provides robust endpoint redundancy, priority, health checking, and failover capabilities.

For an array containing more than two tape libraries, the client should ideally be able to configure multiple alternative endpoints and, where appropriate, apply policies such as endpoint priority, geographical preference, round-robin distribution, or network-performance-based selection.

It would therefore be useful for IBM to document whether client-side endpoint failover is a supported architecture and, if so, which configuration IBM recommends. In particular, recommendations for enterprise applications such as Telestream DIVA would be valuable.

Load balancer or ADC

A second approach is to place a software or hardware load balancer/Application Delivery Controller (ADC) between S3 clients and the Deep Archive systems.

The load balancer could expose a single virtual IP address or hostname and route requests to a healthy Deep Archive system according to policies such as site priority, health status, geographical location, or traffic distribution.

This approach provides centralized control and can implement sophisticated failover policies. It also allows the S3 clients to remain unaware of the underlying multi-site topology.

However, the load balancer becomes an additional component in the S3 data path. It must therefore be deployed redundantly and sized for the expected S3 throughput to avoid introducing either a single point of failure or a throughput bottleneck. The additional network hop and processing overhead should also be considered for high-throughput archive workloads.

IBM has reportedly tested HAProxy in conjunction with a Deep Archive multi-site configuration, routing S3 requests through the S3 virtual IP address of a healthy site. It would be valuable for IBM to clarify whether this is a supported architecture and, if so, provide configuration and health-check recommendations.

Authoritative DNS-based endpoint selection

A third approach is to use an authoritative DNS service to expose a common logical S3 hostname while dynamically selecting the appropriate Deep Archive system.

In this architecture, the DNS service determines which tape-library endpoint should be returned according to availability, priority, geographical topology, or other operational policies. Once DNS resolution has selected a system, the S3 payload traffic flows directly between the client and that Deep Archive system. DNS therefore does not become part of the S3 data path, unlike a load balancer.

This can provide a scalable architecture without introducing an additional component into the high-throughput S3 path. However, it requires a reliable mechanism for determining whether a Deep Archive system is actually capable of serving S3 traffic.

The distinction between simple network reachability and application-level service availability is particularly important here. An ICMP ping or TCP connection may succeed even when the Deep Archive system is not in a state in which it should receive new S3 workloads.

A suitable health check should therefore ideally determine the application-level eligibility of a Deep Archive system to serve S3 requests.

Potential health-check mechanisms include:

  • network-level connectivity checks such as ICMP;

REST APIs exposed by the IBM DiamondBack Control Unit;

REST APIs exposed by the TapeCloud nodes;

an S3-level request such as HeadBucket or HeadObject against a dedicated test bucket or object.

An S3-level health check is particularly interesting because it tests the same service path used by production clients. A lightweight request against a dedicated bucket or object could potentially verify not only network connectivity but also the operational availability of the S3 service.

However, such a mechanism should preferably be explicitly supported and documented by IBM rather than being inferred from the currently exposed APIs.

Requested IBM guidance

The main objective of this idea is therefore to request that IBM document a supported reference architecture for redundant S3 client access to multiple IBM Storage Deep Archive systems, particularly two systems configured for replication.

For example, IBM could clarify:

  1. What reference architecture does IBM recommend for providing redundant S3 access to two Deep Archive systems, with automatic failover from a primary system to a secondary system?

Is client-side endpoint failover supported and recommended? If so, which S3 client configuration or mechanism should be used, and are there specific recommendations for enterprise applications such as Telestream DIVA?

Is the use of an external load balancer or ADC, such as HAProxy or an equivalent enterprise product, supported? If so, what health-check mechanism should be used to determine whether a Deep Archive system is eligible to receive S3 traffic?

Is DNS-based failover supported or recommended? If so, does IBM recommend a particular DNS architecture or product, including open-source solutions such as PowerDNS?

Does IBM provide a native API, agent, plugin, or integration mechanism that can determine the current S3 service eligibility of a Deep Archive system and automatically drive DNS or load-balancer failover?

If an external component must perform the health monitoring, which Deep Archive endpoint or API should be queried? What response or system state defines a Deep Archive instance as “ready” or “eligible to serve S3 traffic”?

What polling interval, timeout, retry, and failure-detection policies does IBM recommend for such health checks?

Does IBM provide a lightweight, non-destructive application-level health check specifically intended for this purpose?

If no native health-check mechanism is provided, does IBM document a supported procedure for implementing one using the existing Deep Archive APIs?

Are there any restrictions or considerations related to replication state, synchronization status, tape-library availability, TapeCloud node state, or ongoing recovery operations that should be taken into account before directing S3 traffic to a particular system?

Why this is important

For enterprise archive environments, S3 endpoint availability is not simply a network-routing problem. The endpoint-selection mechanism must understand whether a particular Deep Archive system is actually able to accept and process S3 workloads.

A documented IBM reference architecture would allow customers and system integrators to implement resilient multi-site S3 access without having to reverse-engineer the operational state of TapeCloud or the Deep Archive platform.

The recommendation could be purely architectural and would not necessarily require changes to the Deep Archive software. For example, IBM could document a supported combination of client-side failover, load balancing, or DNS-based endpoint selection together with an officially supported health-check API.

Such guidance would significantly simplify the design of highly available Deep Archive deployments and provide customers with a clear, supportable approach for implementing S3 failover across replicated tape-library systems.

Idea priority Medium