An apparatus comprises at least one processing device that includes a processor coupled to a memory. The at least one processing device is configured to send commands from a host device to respective ones of a plurality of storage targets of a storage system over a network, to receive in response to the commands, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target, to extract the fault domain identifiers of the respective storage targets from the received storage target data structures, and to store the extracted fault domain identifiers of the respective storage targets in the host device in association with network addresses of the respective storage targets. The extracted fault domain identifiers are illustratively utilized in the host device to generate fault domain connectivity information and/or to perform path selection.
Legal claims defining the scope of protection, as filed with the USPTO.
at least one processing device comprising a processor coupled to a memory; to send commands from a host device to respective ones of a plurality of storage targets of a storage system over a network; to receive in response to the commands, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target; to extract the fault domain identifiers of the respective storage targets from the received storage target data structures; and to store the extracted fault domain identifiers of the respective storage targets in the host device in association with network addresses of the respective storage targets. the at least one processing device being configured: . An apparatus comprising:
claim 1 . The apparatus ofwherein the storage system comprises a distributed storage system that includes a plurality of storage nodes, and further wherein the storage targets are implemented on respective different ones of the storage nodes.
claim 1 . The apparatus ofthe storage targets each comprise one or more Non-Volatile Memory Express (NVMe) controllers of the storage system.
claim 1 . The apparatus ofwherein the commands comprise respective NVMe Identify commands and the storage target data structures comprise respective NVMe Identify Controller data structures.
claim 1 . The apparatus ofwherein at least a portion of a given one of the fault domain identifiers is extracted from at least one reserved field of the corresponding storage target data structure.
claim 1 . The apparatus ofwherein at least a portion of a given one of the fault domain identifiers is extracted from at least one vendor-specific field of the corresponding storage target data structure.
claim 1 to generate fault domain connectivity information based at least in part on the extracted fault domain identifiers, the fault domain connectivity information indicating different types of network connectivity between the host device and respective fault domains identified by the extracted fault domain identifiers; and to perform one or more automated actions based at least in part on the generated fault domain connectivity information. . The apparatus ofwherein the at least one processing device is further configured:
claim 7 . The apparatus ofwherein the one or more automated actions comprise controlling initiation of a replication process by the host device.
claim 7 . The apparatus ofwherein the one or more automated actions comprise generating an alert specifying an absence of a particular type of network connectivity to at least one of the fault domains.
claim 7 . The apparatus ofwherein the different types of network connectivity comprise at least full connectivity, partial connectivity and no connectivity.
claim 1 at least one storage node of a plurality of storage nodes of the storage system; at least one physical server of a plurality of physical servers of the storage system; and at least one subnetwork of a plurality of subnetworks of the storage system. . The apparatus ofwherein a given fault domain identified by a corresponding one of the extracted fault domain identifiers comprises at least one of:
claim 1 . The apparatus ofwherein the at least one processing device is further configured to control path selection for delivery of input-output operations from the host device to the storage targets based at least in part on the extracted fault domain identifiers.
claim 12 avoiding selection of multiple paths to respective storage targets that have the same fault domain identifier; preventing selection of multiple paths to respective storage targets that have the same fault domain identifier; and load balancing delivery of a plurality of input-output operations over respective ones of a plurality of paths to respective ones of multiple storage targets each having a different fault domain identifier in a manner that ensures that each of the plurality of input-output operations is delivered over a path to a storage target having a different fault domain identifier. . The apparatus ofwherein controlling path selection comprises at least one of:
claim 1 . The apparatus ofwherein the extracted fault domain identifiers are utilized by the host device to match multiple network addresses for respective different storage targets to a same fault domain of the storage system.
to send commands from a host device to respective ones of a plurality of storage targets of a storage system over a network; to receive in response to the commands, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target; to extract the fault domain identifiers of the respective storage targets from the received storage target data structures; and to store the extracted fault domain identifiers of the respective storage targets in the host device in association with network addresses of the respective storage targets. . A computer program product comprising a non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device comprising a processor coupled to a memory, causes the at least one processing device:
claim 15 to generate fault domain connectivity information based at least in part on the extracted fault domain identifiers, the fault domain connectivity information indicating different types of network connectivity between the host device and respective fault domains identified by the extracted fault domain identifiers; and to perform one or more automated actions based at least in part on the generated fault domain connectivity information. . The computer program product ofwherein the program code when executed by the at least one processing device further causes the at least one processing device:
claim 15 . The computer program product ofwherein the program code when executed by the at least one processing device further causes the at least one processing device to control path selection for delivery of input-output operations from the host device to the storage targets based at least in part on the extracted fault domain identifiers.
sending commands from a host device to respective ones of a plurality of storage targets of a storage system over a network; receiving in response to the commands, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target; extracting the fault domain identifiers of the respective storage targets from the received storage target data structures; and storing the extracted fault domain identifiers of the respective storage targets in the host device in association with network addresses of the respective storage targets; wherein the method is performed by at least one processing device comprising a processor coupled to a memory. . A method comprising:
claim 18 generating fault domain connectivity information based at least in part on the extracted fault domain identifiers, the fault domain connectivity information indicating different types of network connectivity between the host device and respective fault domains identified by the extracted fault domain identifiers; and performing one or more automated actions based at least in part on the generated fault domain connectivity information. . method ofwherein the method further comprises:
claim 18 . The method ofwherein the method further comprises controlling path selection for delivery of input-output operations from the host device to the storage targets based at least in part on the extracted fault domain identifiers.
Complete technical specification and implementation details from the patent document.
The field relates generally to information processing systems, and more particularly to storage in information processing systems.
Information processing systems often include distributed storage systems comprising multiple storage nodes. These distributed storage systems may be dynamically reconfigurable under software control in order to adapt the number and type of storage nodes and the corresponding system storage capacity as needed, in an arrangement commonly referred to as a software-defined storage system. For example, in a typical software-defined storage system, storage capacities of multiple distributed storage nodes are pooled together into one or more storage pools. For applications running on a host that utilizes the software-defined storage system, such a storage system provides a logical storage object view to allow a given application to store and access data, without the application being aware that the data is being dynamically distributed among different storage nodes. In these and other distributed storage system arrangements, it can be unduly difficult to discover adequate information regarding targets of respective storage nodes, particularly when using advanced storage access protocols such as Non-Volatile Memory Express (NVMe) over Fabrics, also referred to as NVMe-oF, or NVMe over Transmission Control Protocol (TCP), also referred to as NVMe/TCP. For example, conventional approaches can lead to sub-optimal arrangements in terms of resiliency and load balancing, thereby adversely impacting storage system performance.
Illustrative embodiments disclosed herein provide techniques for host device processing of fault domain identifiers obtained from respective storage targets in a distributed storage system or other type of storage system. Such techniques can advantageously facilitate host-side functionality such as generating fault domain connectivity information and/or performing path selection, particularly when utilizing advanced storage access protocols such as NVMe-oF and NVMe/TCP, as well as in other storage contexts.
For example, some embodiments utilize the disclosed fault domain identifier processing techniques to provide enhanced resiliency and load balancing in delivery of input-output (IO) operations from one or more host devices to a storage array or other storage system over paths through a network. These and other embodiments can significantly improve storage system performance.
In one embodiment, an apparatus comprises at least one processing device that includes a processor coupled to a memory. The at least one processing device is configured to send commands from a host device to respective ones of a plurality of storage targets of a storage system over a network, to receive in response to the commands, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target, to extract the fault domain identifiers of the respective storage targets from the received storage target data structures, and to store the extracted fault domain identifiers of the respective storage targets in the host device in association with network addresses of the respective storage targets.
As indicated above, in some embodiments, the extracted fault domain identifiers are illustratively utilized in the host device to generate fault domain connectivity information and/or to perform path selection. Additional or alternative fault domain processing functionality can be implemented using the extracted fault domain identifiers in other embodiments.
In some embodiments, the storage system comprises a distributed storage system that includes a plurality of storage nodes, with the storage targets being implemented on respective different ones of the storage nodes. In other embodiments, the disclosed techniques can be implemented using stand-alone storage arrays or other types of storage systems that are not distributed across multiple storage nodes. Accordingly, the disclosed techniques are applicable to a wide variety of different types of storage systems.
The storage targets in some embodiments each comprise one or more NVMe controllers of the storage system, although other types of storage targets can be used in other embodiments.
A given fault domain identified by a corresponding one of the extracted fault domain identifiers in some embodiments illustratively comprises at least one storage node of a plurality of storage nodes of the storage system, at least one physical server of a plurality of physical servers of the storage system, and/or at least one subnetwork of a plurality of subnetworks of the storage system. Other types of fault domains may be used in other embodiments.
In some embodiments, the commands sent from the host device to respective ones of the storage targets comprise respective NVMe Identify commands, and the storage target data structures comprise respective NVMe Identify Controller data structures, each illustratively modified as disclosed herein to include a fault domain identifier for the corresponding storage target, although other types of commands and storage target data structures may be used in other embodiments.
At least a portion of a given one of the fault domain identifiers in some embodiments is illustratively extracted from at least one reserved field of the corresponding storage target data structure, such as one or more reserved fields of the above-noted NVMe Identify Controller data structure.
Additionally or alternatively, at least a portion of a given one of the fault domain identifiers in some embodiments is illustratively extracted from at least one vendor-specific field of the corresponding storage target data structure, such as one or more vendor-specific fields of the above-noted NVMe Identify Controller data structure.
In some embodiments, the at least one processing device is further configured to generate fault domain connectivity information based at least in part on the extracted fault domain identifiers, with the fault domain connectivity information indicating different types of network connectivity between the host device and respective fault domains identified by the extracted fault domain identifiers, and to perform one or more automated actions based at least in part on the generated fault domain connectivity information. The one or more automated actions may comprise, for example, controlling initiation of a replication process by the host device and/or generating an alert specifying an absence of a particular type of network connectivity to at least one of the fault domains.
Additionally or alternatively, the at least one processing device in some embodiments is further configured to control path selection for delivery of IO operations from the host device to the storage targets based at least in part on the extracted fault domain identifiers. Controlling path selection can comprise, for example, avoiding selection of multiple paths to respective storage targets that have the same fault domain identifier, preventing selection of multiple paths to respective storage targets that have the same fault domain identifier, and/or load balancing delivery of a plurality of IO operations over respective ones of a plurality of paths to respective ones of multiple storage targets each having a different fault domain identifier in a manner that ensures that each of the plurality of IO operations is delivered over a path to a storage target having a different fault domain identifier.
Such features of illustrative embodiments are examples only, and should not be viewed as limiting in any way.
These and other illustrative embodiments include, without limitation, apparatus, systems, methods and computer program products comprising processor-readable storage media.
Illustrative embodiments will be described herein with reference to exemplary information processing systems and associated computers, servers, storage devices and other processing devices. It is to be appreciated, however, that these and other embodiments are not restricted to the particular illustrative system and device configurations shown. Accordingly, the term “information processing system” as used herein is intended to be broadly construed, so as to encompass, for example, processing systems comprising cloud computing and storage systems, as well as other types of processing systems comprising various combinations of physical and virtual processing resources. An information processing system may therefore comprise, for example, at least one data center or other cloud-based system that includes one or more clouds hosting multiple tenants that share cloud resources, as well as other types of systems comprising a combination of cloud and edge infrastructure. Numerous different types of enterprise computing and storage systems are also encompassed by the term “information processing system” as that term is broadly used herein.
1 FIG. 100 100 101 1 101 2 101 101 102 101 101 102 104 104 104 shows an information processing systemconfigured in accordance with an illustrative embodiment. The information processing systemcomprises a plurality of hosts-,-, . . .-N, collectively referred to herein as hosts, and a distributed storage systemshared by the hosts. The hostsand distributed storage systemin this embodiment are configured to communicate with one another via a networkthat illustratively utilizes protocols such as Transmission Control Protocol (TCP) and Internet Protocol (IP), and is therefore referred to herein as a TCP/IP network, although it is to be appreciated that the networkcan operate using additional or alternative protocols. In some embodiments, the networkcomprises a storage area network (SAN) that includes one or more Fibre Channel (FC) switches, Ethernet switches or other types of switch fabrics.
101 101 The hostsare also referred to herein as respective “host devices.” It should be noted that terms such as “host” and “host device” as used herein are intended to be broadly construed, so as to encompass, for example, a host system which may comprise multiple distinct devices of various types. A given one of the hostsin some embodiments can therefore comprise, for example, at least one server, as well as a wide variety of additional or alternative types and arrangements of processing devices.
102 105 1 105 2 105 105 The distributed storage systemmore particularly comprises a plurality of storage nodes-,-, . . .-M, collectively referred to herein as storage nodes. The values N and M in this embodiment denote arbitrary integer values that in the figure are illustrated as being greater than or equal to three, although other values such as N=1, N=2, M=1 or M=2 can be used in other embodiments.
105 102 The storage nodescollectively form the distributed storage system, which is just one possible example of what is generally referred to herein as a “distributed storage system.” Other distributed storage systems can include different numbers and arrangements of storage nodes, and possibly one or more additional components. For example, as indicated above, a distributed storage system in some embodiments may include only first and second storage nodes, corresponding to an M=2 embodiment. Some embodiments can configure a distributed storage system to include additional components in the form of a system manager implemented using one or more additional nodes.
102 105 105 105 105 105 105 In some embodiments, the distributed storage systemprovides a logical address space that is divided among the storage nodes, such that different ones of the storage nodesstore the data for respective different portions of the logical address space. Accordingly, in these and other similar distributed storage system arrangements, different ones of the storage nodeshave responsibility for different portions of the logical address space. For a given logical storage volume, logical blocks of that logical storage volume are illustratively distributed across the storage nodes. Additionally or alternatively, logical blocks of one or more logical storage volumes may each be accessible via only a subset of the storage nodes. For example, a given one of the storage nodesmay store an entire logical storage volume, or multiple entire logical storage volumes.
102 105 105 Other types of distributed storage systems can be used in other embodiments. For example, distributed storage systemcan comprise multiple distinct storage arrays, such as a production storage array and a backup storage array, possibly deployed at different locations. Accordingly, in some embodiments, one or more of the storage nodesmay each be viewed as comprising at least a portion of a separate storage array with its own logical address space. Alternatively, the storage nodescan be viewed as collectively comprising one or more storage arrays. The term “storage node” as used herein is therefore intended to be broadly construed.
102 105 105 3 FIG. In some embodiments, the distributed storage systemcomprises a software-defined storage system and the storage nodescomprise respective software-defined storage server nodes of the software-defined storage system, such nodes also being referred to herein as SDS server nodes, where SDS denotes software-defined storage. Accordingly, the number and types of storage nodescan be dynamically expanded or contracted under software control in some embodiments. Examples of such software-defined storage systems will be described in more detail below in conjunction with.
102 It is to be appreciated, however, that techniques disclosed herein can be implemented in other embodiments in stand-alone storage arrays or other types of storage systems that are not distributed across multiple storage nodes. The disclosed techniques are therefore applicable to a wide variety of different types of storage systems. The distributed storage systemis just one illustrative example.
102 105 101 101 In the distributed storage system, each of the storage nodesis illustratively configured to interact with one or more of the hosts. The hostsillustratively comprise servers or other types of computers of an enterprise computer system, cloud-based computer system or other arrangement of multiple compute nodes, each associated with one or more system users.
101 101 105 105 The hostsin some embodiments illustratively provide compute services such as execution of one or more applications on behalf of each of one or more users associated with respective ones of the hosts. Such applications illustratively generate input-output (IO) operations that are processed by a corresponding one of the storage nodes. The term “input-output” as used herein refers to at least one of input and output. For example, IO operations may comprise write requests and/or read requests directed to logical addresses of a particular logical storage volume of one or more of the storage nodes. These and other types of IO operations are also generally referred to herein as IO requests.
102 105 100 105 101 105 102 101 The IO operations that are currently being processed in the distributed storage systemin some embodiments are referred to herein as outstanding IOs that have been admitted by the storage nodesto further processing within the system. The storage nodesare illustratively configured to queue IO operations arriving from one or more of the hostsin one or more sets of storage-side IO queues, to be distinguished from host-side IO queues referred to elsewhere herein. In some embodiments, each of the storage nodescomprises one or more NVMe targets or other types of storage targets of the distributed storage system, also referred to herein as simply “targets,” and each such target is illustratively configured with a plurality of storage-side IO queues. Each such storage-side IO queue or set of storage-side IO queues may have at least one corresponding TCP connection or other type of network connection with one or more of the hosts. Such IO queues and network connections are considered examples of “network resources” as that term is broadly used herein.
105 105 The storage nodesillustratively comprise respective processing devices of one or more processing platforms. For example, the storage nodescan each comprise one or more processing devices each having a processor and a memory, possibly implementing virtual machines and/or containers, although numerous other configurations are possible.
105 The storage nodescan additionally or alternatively be part of cloud infrastructure, such as a cloud-based system implementing Storage-as-a-Service (STaaS) functionality.
105 The storage nodesmay be implemented on a common processing platform, or on separate processing platforms. In the case of separate processing platforms, there may be a single storage node per processing platform or multiple storage nodes per processing platform.
101 102 105 101 The hostsare illustratively configured to write data to and read data from the distributed storage systemcomprising storage nodesin accordance with applications executing on those hostsfor system users.
The term “user” herein is intended to be broadly construed so as to encompass numerous arrangements of human, hardware, software or firmware entities, as well as combinations of such entities. Compute and/or storage services may be provided for users under a Platform-as-a-Service (PaaS) model, an Infrastructure-as-a-Service (IaaS) model and/or a Function-as-a-Service (FaaS) model, although it is to be appreciated that numerous other cloud infrastructure arrangements could be used. Also, illustrative embodiments can be implemented outside of the cloud infrastructure context, as in the case of a stand-alone computing and storage system implemented within a given enterprise. Combinations of cloud and edge infrastructure can also be used in implementing a given information processing system to provide services to users.
100 100 104 Communications between the components of systemcan take place over additional or alternative networks, including a global computer network such as the Internet, a wide area network (WAN), a local area network (LAN), a satellite network, a telephone or cable network, a cellular network such as 4G or 5G cellular network, a wireless network such as a WiFi or WiMAX network, or various portions or combinations of these and other types of networks. The systemin some embodiments therefore comprises one or more additional networks other than networkeach comprising processing devices configured to communicate using TCP, IP and/or other communication protocols.
As a more particular example, some embodiments may utilize one or more high-speed local networks in which associated processing devices communicate with one another utilizing Peripheral Component Interconnect express (PCIe) interface cards of those devices, that support networking protocols such as InfiniBand or Fibre Channel, in addition to or in place of TCP/IP. Numerous alternative networking arrangements are possible in a given embodiment, as will be appreciated by those skilled in the art. Additional examples include remote direct memory access (RDMA) over Converged Ethernet (RoCE) or RDMA over iWARP.
105 1 106 1 108 1 106 1 102 106 1 105 1 105 1 105 2 105 105 The first storage node-comprises a plurality of storage devices-and an associated storage processor-. The storage devices-illustratively store metadata pages and user data pages associated with one or more storage volumes of the distributed storage system. The storage volumes illustratively comprise respective logical units (LUNs) or other types of logical storage volumes (e.g., NVMe namespaces). The storage devices-in some embodiments more particularly comprise local persistent storage devices of the first storage node-. Such persistent storage devices are local to the first storage node-, but remote from the second storage node-, the storage node-M and any other ones of other storage nodes.
105 2 105 105 1 105 2 106 2 108 2 105 106 108 Each of the other storage nodes-through-M is assumed to be configured in a manner similar to that described above for the first storage node-. Accordingly, by way of example, storage node-comprises a plurality of storage devices-and an associated storage processor-, and storage node-M comprises a plurality of storage devices-M and an associated storage processor-M.
106 2 106 102 106 2 105 2 105 2 105 1 105 105 106 105 105 105 1 105 2 105 As indicated previously, the storage devices-through-M illustratively store metadata pages and user data pages associated with one or more storage volumes of the distributed storage system, such as the above-noted LUNs or other types of logical storage volumes. The storage devices-in some embodiments more particularly comprise local persistent storage devices of the storage node-. Such persistent storage devices are local to the storage node-, but remote from the first storage node-, the storage node-M, and any other ones of the storage nodes. Similarly, the storage devices-M in some embodiments more particularly comprise local persistent storage devices of the storage node-M. Such persistent storage devices are local to the storage node-M, but remote from the first storage node-, the second storage node-, and any other ones of the storage nodes.
105 The local persistent storage of a given one of the storage nodesillustratively comprises the particular local persistent storage devices that are implemented in or otherwise associated with that storage node.
1 FIG. 102 108 106 In theembodiment, the distributed storage systemcomprises storage processorsand corresponding sets of storage devices, and may include additional or alternative components, such as sets of local caches.
108 102 101 108 101 106 108 108 The storage processorsillustratively control the processing of IO operations received in the distributed storage systemfrom the hosts. For example, the storage processorsillustratively manage the processing of read and write commands directed by the MPIO drivers of the hoststo particular ones of the storage devices. The storage processorscan be implemented as respective storage controllers, directors or other storage system components configured to control storage system operations relating to processing of IO operations. In some embodiments, each of the storage processorshas a different one of the above-noted local caches associated therewith, although numerous alternative arrangements are possible.
108 105 The storage processorsof the storage nodesmay include additional modules and other components typically found in conventional implementations of storage processors and storage systems, although such additional modules and other components are omitted from the figure for clarity and simplicity of illustration.
108 105 Additionally or alternatively, the storage processorsin some embodiments can comprise or be otherwise associated with one or more write caches and one or more write cache journals, both also illustratively distributed across the storage nodesof the distributed storage system. It is further assumed in illustrative embodiments that one or more additional journals are provided in the distributed storage system, such as, for example, a metadata update journal and possibly other journals providing other types of journaling functionality for IO operations. Illustrative embodiments disclosed herein are assumed to be configured to perform various destaging processes for write caches and associated journals, and to perform additional or alternative functions in conjunction with processing of IO operations.
106 105 106 The storage devicesof the storage nodesillustratively comprise solid state drives (SSDs). Such SSDs are implemented using non-volatile memory (NVM) devices such as flash memory. Other types of NVM devices that can be used to implement at least a portion of the storage devicesinclude, for example, non-volatile random access memory (NVRAM), phase-change RAM (PC-RAM), magnetic RAM (MRAM), resistive RAM, and spin torque transfer magneto-resistive RAM (STT-MRAM). These and various combinations of multiple different types of NVM devices may also be used. For example, hard disk drives (HDDs) can be used in combination with or in place of SSDs or other types of NVM devices.
106 105 1 FIG. However, it is to be appreciated that other types of storage devices can be used in other embodiments. For example, a given storage system as the term is broadly used herein can include a combination of different types of storage devices, as in the case of a multi-tier storage system comprising a flash-based fast tier and a disk-based capacity tier. In such an embodiment, each of the fast tier and the capacity tier of the multi-tier storage system comprises a plurality of storage devices with different types of storage devices being used in different ones of the storage tiers. For example, the fast tier may comprise flash drives while the capacity tier comprises HDDs. The particular storage devices used in a given storage tier may be varied in other embodiments, and multiple distinct storage device types may be used within a single storage tier. The term “storage device” as used herein is intended to be broadly construed, so as to encompass, for example, SSDs, HDDs, flash drives, hybrid drives or other types of storage devices. Such storage devices are examples of local persistent storage devices that may be used to implement at least a portion of the storage devicesof the storage nodesof the distributed storage system of.
105 105 In some embodiments, the storage nodescollectively provide a distributed storage system, although the storage nodescan be used to implement other types of storage systems in other embodiments. One or more such storage nodes can be associated with at least one storage array. Additional or alternative types of storage products that can be used in implementing a given storage system in illustrative embodiments include software-defined storage, cloud storage and object-based storage. Combinations of multiple ones of these and other storage types can also be used.
105 105 As indicated above, the storage nodesin some embodiments comprise respective software-defined storage server nodes of a software-defined storage system, in which the number and types of storage nodescan be dynamically expanded or contracted under software control using software-defined storage techniques.
The term “storage system” as used herein is therefore intended to be broadly construed, and should not be viewed as being limited to certain types of storage systems, such as content addressable storage systems or flash-based storage systems. A given storage system as the term is broadly used herein can comprise, for example, network-attached storage (NAS), storage area networks (SANs), direct-attached storage (DAS) and distributed DAS, as well as combinations of these and other storage types, including software-defined storage.
101 105 In some embodiments, communications between the hostsand the storage nodescomprise NVMe commands of an NVMe storage access protocol, for example, as described in the NVM Express Base Specification, Revision 2.1, August 2024, and its associated NVM Express Command Set Specification and NVM Express TCP Transport Specification, all of which are incorporated by reference herein. Other examples of NVMe storage access protocols that may be utilized in illustrative embodiments disclosed herein include NVMe over Fabrics, also referred to herein as NVMe-oF, and NVMe over TCP, also referred to herein as NVMe/TCP. Other embodiments can utilize other types of storage access protocols, including, for example, NVMe over Fibre Channel, also referred to herein as NVMe/FC.
101 105 As another example, communications between the hostsand the storage nodesin some embodiments can be implemented using Small Computer System Interface (SCSI) commands and the Internet SCSI (iSCSI) protocol.
Other types of commands may be used in other embodiments, including commands that are part of a standard command set, or custom commands such as a “vendor unique command” or VU command that is not part of a standard command set. The term “command” as used herein is therefore intended to be broadly construed, so as to encompass, for example, a composite command that comprises a combination of multiple individual commands. Numerous other types, formats and configurations of IO operations can be used in other embodiments, as that term is broadly used herein.
106 105 102 Some embodiments disclosed herein are configured to utilize one or more RAID arrangements to store data across the storage devicesin each of one or more of the storage nodesof the distributed storage system. Other embodiments can utilize other data protection techniques, such as, for example, Erasure Coding (EC), instead of one or more RAID arrangements.
The RAID arrangement can comprise, for example, a RAID 5 arrangement supporting recovery from a failure of a single one of the plurality of storage devices, a RAID 6 arrangement supporting recovery from simultaneous failure of up to two of the storage devices, or another type of RAID arrangement. For example, some embodiments can utilize RAID arrangements with redundancy higher than two.
The term “RAID arrangement” as used herein is intended to be broadly construed, and should not be viewed as limited to RAID 5, RAID 6 or other parity RAID arrangements. For example, a RAID arrangement in some embodiments can comprise combinations of multiple instances of distinct RAID approaches, such as a mixture of multiple distinct RAID types (e.g., RAID 1 and RAID 6) over the same set of storage devices, or a mixture of multiple stripe sets of different instances of one RAID type (e.g., two separate instances of RAID 5) over the same set of storage devices. Other types of parity RAID techniques and/or non-parity RAID techniques can be used in other embodiments.
108 105 106 Such a RAID arrangement is illustratively established by the storage processorsof the respective storage nodes. The storage devicesin the context of RAID arrangements herein are also referred to as “disks” or “drives.” A given such RAID arrangement may also be referred to in some embodiments herein as a “RAID array.”
106 The RAID arrangement used in an illustrative embodiment includes a plurality of devices, each illustratively a different physical storage device of the storage devices. Multiple such physical storage devices are typically utilized to store data of a given LUN or other logical storage volume in the distributed storage system. For example, data pages or other data blocks of a given LUN or other logical storage volume can be “striped” along with its corresponding parity information across multiple ones of the devices in the RAID arrangement in accordance with RAID 5 or RAID 6 techniques.
A given RAID 5 arrangement defines block-level striping with single distributed parity and provides fault tolerance of a single drive failure, so that the array continues to operate with a single failed drive, irrespective of which drive fails. For example, in a conventional RAID 5 arrangement, each stripe includes multiple data blocks as well as a corresponding p parity block. The p parity blocks are associated with respective row parity information computed using well-known RAID 5 techniques. The data and parity blocks are distributed over the devices to support the above-noted single distributed parity and its associated fault tolerance.
A given RAID 6 arrangement defines block-level striping with double distributed parity and provides fault tolerance of up to two drive failures, so that the array continues to operate with up to two failed drives, irrespective of which two drives fail. For example, in a conventional RAID 6 arrangement, each stripe includes multiple data blocks as well as corresponding p and q parity blocks. The p and q parity blocks are associated with respective row parity information and diagonal parity information computed using well-known RAID 6 techniques. The data and parity blocks are distributed over the devices to collectively provide a diagonal-based configuration for the p and q parity information, so as to support the above-noted double distributed parity and its associated fault tolerance.
In such RAID arrangements, the parity blocks are typically not read unless needed for a rebuild process triggered by one or more storage device failures.
106 105 These and other references herein to RAID 5, RAID 6 and other particular RAID arrangements are only examples, and numerous other RAID arrangements can be used in other embodiments. Also, other embodiments can store data across the storage devicesof the storage nodeswithout using RAID arrangements.
105 106 105 105 106 1 FIG. In some embodiments, the storage nodesof the distributed storage system ofare connected to each other in a full mesh network, and are collectively managed by a system manager. A given set of local persistent storage devices or other storage deviceson a given one of the storage nodesis illustratively implemented in a disk array enclosure (DAE) or other type of storage array enclosure of that storage node. Each of the storage nodesillustratively comprises a CPU or other type of processor, a memory, a network interface card (NIC) or other type of network interface, and its corresponding storage devices, possibly arranged as part of a DAE of the storage node.
105 105 In some embodiments, different ones of the storage nodesare associated with the same DAE or other type of storage array enclosure. The system manager is illustratively implemented as a management module or other similar management logic instance, possibly running on one or more of the storage nodes, on another storage node and/or on a separate non-storage node of the distributed storage system.
105 As a more particular non-limiting illustration, two or more of the storage nodesin some embodiments are arranged together in multiple groups, with each such group being coupled to a different DAE comprising multiple drives, and each node in a group being connected to the DAE and to each drive through a separate connection. The system manager may be running on one of the nodes of a first one of the groups of the distributed storage system. Again, numerous other arrangements of the storage nodes are possible in a given distributed storage system as disclosed herein.
100 110 112 116 110 112 105 105 112 110 105 The systemas shown further comprises a plurality of system management nodesthat are illustratively configured to provide system management functionality of the type noted above. Such functionality in the present embodiment illustratively further involves utilization of control plane serversand a system management database. In some embodiments, at least portions of the system management nodesand their associated control plane serversare distributed over the storage nodes. For example, a designated subset of the storage nodescan each be configured to include a corresponding one of the control plane servers. Other system management functionality provided by system management nodescan be similarly distributed over a subset of the storage nodes.
116 100 The system management databasestores configuration and operation information of the systemand portions thereof are illustratively accessible to various system administrators such as host administrators and storage administrators.
101 1 101 2 101 114 1 114 2 114 114 102 108 105 The hosts-,-, . . .-N include respective instances of path selection logic-,-, . . .-N. Such instances of path selection logicare illustratively utilized in supporting functionality for fault domain identifier processing in the distributed storage system, illustratively through interaction with fault domain identifier processing logic instances implemented in respective ones of the storage processorsof the storage nodes, as described in more detail below.
105 102 101 In some embodiments, each of the storage nodesof the distributed storage systemis assumed to comprise multiple controllers associated with a corresponding target of that storage node. Such a “storage target” or simply “target” as those terms are broadly used herein is illustratively a destination end of one or more paths from one or more of the hoststo the storage node, and may comprise, for example, an NVMe subsystem of the storage node, although other types of targets can be used in other embodiments. It should be noted that different types of targets may be present in NVMe embodiments than are present in other embodiments that use other storage access protocols, such as SCSI embodiments. Accordingly, the types of targets that may be implemented in a given embodiment can vary depending upon the particular storage access protocol being utilized in that embodiment, and/or other factors. Similarly, the types of initiators can vary depending upon the particular storage access protocol, and/or other factors. Again, terms such as “initiator” and “target” as used herein are intended to be broadly construed, and should not be viewed as being limited in any way to particular types of components associated with any particular storage access protocol.
114 101 101 102 The paths that are selected by instances of path selection logicof the hostsfor delivering IO operations from the hoststo the distributed storage systemare associated with respective initiator-target pairs, as described in more detail elsewhere herein.
101 114 101 105 102 102 In some embodiments, IO operations are processed in the hostsutilizing their respective instances of path selection logicin the following manner. A given one of the hostsestablishes a plurality of paths between at least one initiator of the given host and a plurality of targets of respective storage nodesof the distributed storage system. For each of a plurality of IO operations generated in the given host for delivery to the distributed storage system, the host selects a path to a particular target, and sends the IO operation to the corresponding storage node over the selected path.
105 102 The given host above is an example of what is more generally referred to herein as “at least one processing device” that includes a processor coupled to a memory. The storage nodesof the distributed storage systemare also examples of “at least one processing device” as that term is broadly used herein.
101 114 It is to be appreciated that path selection as disclosed herein can be performed independently by each of the hosts, illustratively utilizing their respective instances of path selection logic, as indicated above, with possible involvement of additional or alternative system components.
105 In some embodiments, the initiator of the given host and the targets of the respective storage nodesare configured to support one or more designated standard storage access protocols, such as an NVMe access protocol or a SCSI access protocol. As more particular examples in the NVMe context, the designated storage access protocol utilized in some embodiments may comprise an NVMe-oF, NVMe/TCP or NVMe/FC access protocol, although a wide variety of additional or alternative storage access protocols can be used in other embodiments.
101 101 101 101 102 114 114 101 The hostscan comprise additional or alternative components. For example, in some embodiments, the hostsfurther comprise respective sets of host-side IO queues and respective multi-path input-output (MPIO) drivers. The MPIO drivers collectively comprise a multi-path layer of the hosts. Path selection functionality for delivery of IO operations from the hoststo the distributed storage systemis provided in the multi-path layer by respective instances of path selection logicimplemented within the MPIO drivers. In some embodiments, the instances of path selection logicare implemented at least in part within the MPIO drivers of the hosts.
The MPIO drivers may comprise, for example, otherwise conventional MPIO drivers, such as PowerPath® drivers from Dell Technologies, suitably modified in the manner disclosed herein to provide one or more portions of the disclosed functionality for fault domain identifier processing. Other types of MPIO drivers from other driver vendors may be suitably modified to incorporate one or more portions of the functionality for fault domain identifier processing as disclosed herein.
114 101 For example, the instances of path selection logicof the respective hostscan be implemented at least in part in respective MPIO drivers of those hosts.
114 102 In some embodiments, such instances of path selection logicinclude or are otherwise associated with respective corresponding instances of host-side fault domain identifier processing logic that are configured to send commands from the corresponding host to respective ones of a plurality of storage targets of the distributed storage system, to receive in response to the commands, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target, to extract the fault domain identifiers of the respective storage targets from the received storage target data structures, and to store the extracted fault domain identifiers of the respective storage targets in the host in association with network addresses of the respective storage targets. The stored fault domain identifiers are illustratively utilized by the host to implement additional host-side functionality such as generating fault domain connectivity information and/or performing path selection, as will be described in more detail elsewhere herein.
101 114 101 The above-described host-side fault domain identifier processing logic is illustratively implemented at least in part within an MPIO layer of the hosts. For example, it may be deployed at least in part within a corresponding instance of path selection logicin the MPIO layer, or can be implemented elsewhere within the hosts.
101 101 In some embodiments, the hostscomprise respective local caches, implemented using respective memories of those hosts. A given such local cache can be implemented using one or more cache cards. A wide variety of different caching techniques can be used in other embodiments, as will be appreciated by those skilled in the art. Other examples of memories of the respective hoststhat may be utilized to provide local caches include one or more memory cards or other memory devices, such as, for example, an NVMe over PCIe cache card, a local flash drive or other type of NVM storage drive, or combinations of these and other host memory devices.
102 104 101 101 102 104 100 The MPIO drivers are illustratively configured to deliver IO operations selected from their respective sets of host-side IO queues to the distributed storage systemvia selected ones of multiple paths over the network. The sources of the IO operations stored in the sets of host-side IO queues illustratively include respective processes of one or more applications executing on the hosts. For example, IO operations can be generated by each of multiple processes of a database application running on one or more of the hosts. Such processes issue IO operations for delivery to the distributed storage systemover the network. Other types of sources of IO operations may be present in a given implementation of system.
101 A given IO operation is therefore illustratively generated by a process of an application running on a given one of the hosts, and is queued in one of the IO queues of the given host with other operations generated by other processes of that application, and possibly other processes of other applications.
102 106 102 106 The paths from the given host to the distributed storage systemillustratively comprise paths associated with respective initiator-target pairs, with each initiator comprising, for example, a port of a single-port or multi-port host bus adaptor (HBA) or other initiating entity of the given host and each target comprising a port or other targeted entity corresponding to one or more of the storage devicesof the distributed storage system. As noted above, the storage devicesillustratively comprise LUNs or other types of logical storage devices.
102 104 In some embodiments, the paths are associated with respective communication links between the given host and the distributed storage systemwith each such communication link having a negotiated link speed. For example, in conjunction with registration of a given HBA to a switch of the network, the HBA and the switch may negotiate a link speed. The actual link speed that can be achieved in practice in some cases is less than the negotiated link speed, which is a theoretical maximum value.
Negotiated rates of the respective particular initiator and the corresponding target illustratively comprise respective negotiated data rates determined by execution of at least one link negotiation protocol for an associated one of the paths.
In some embodiments, at least a portion of the initiators comprise virtual initiators, such as, for example, respective ones of a plurality of N-Port ID Virtualization (NPIV) initiators associated with one or more Fibre Channel (FC) network connections. Such initiators illustratively utilize NVMe arrangements such as NVMe/FC, although other protocols can be used. Other embodiments can utilize other types of virtual initiators in which multiple network addresses can be supported by a single network interface, such as, for example, multiple media access control (MAC) addresses on a single network interface of an Ethernet network interface card (NIC). Accordingly, in some embodiments, the multiple virtual initiators are identified by respective ones of a plurality of media MAC addresses of a single network interface of a NIC. Such initiators illustratively utilize NVMe arrangements such as NVMe/TCP, although again other protocols can be used.
101 Accordingly, in some embodiments, multiple virtual initiators are associated with a single HBA of a given one of the hostsbut have respective unique identifiers associated therewith.
Additionally or alternatively, different ones of the multiple virtual initiators are illustratively associated with respective different ones of a plurality of virtual machines of the given host that share a single HBA of the given host, or a plurality of logical partitions of the given host that share a single HBA of the given host.
Numerous alternative virtual initiator arrangements are possible, as will be apparent to those skilled in the art. The term “virtual initiator” as used herein is therefore intended to be broadly construed. It is also to be appreciated that other embodiments need not utilize any virtual initiators. References herein to the term “initiators” are intended to be broadly construed, and should therefore be understood to encompass physical initiators, virtual initiators, or combinations of both physical and virtual initiators.
102 104 102 102 Various scheduling algorithms, load balancing algorithms and/or other types of algorithms can be utilized by the MPIO driver of the given host in delivering IO operations from the IO queues of that host to the distributed storage systemover particular paths via the network. Each such IO operation is assumed to comprise one or more commands for instructing the distributed storage systemto perform particular types of storage-related functions such as reading data from or writing data to particular logical volumes of the distributed storage system. Such commands are assumed to have various payload sizes associated therewith, and the payload associated with a given command is referred to herein as its “command payload.”
102 A command directed by the given host to the distributed storage systemis considered an “outstanding” command until such time as its execution is completed in the viewpoint of the given host, at which time it is considered a “completed” command. The commands illustratively comprise respective NVMe commands, although other command formats, such as SCSI command formats, can be used in other embodiments. Command formats such as Submission Queue Entry (SQE) are utilized in the NVMe context. In the SCSI context, a given such command is illustratively defined by a corresponding command descriptor block (CDB) or similar format construct. The given command can have multiple blocks of payload associated therewith, such as a particular number of 512-byte SCSI blocks or other types of blocks.
102 5 FIG. In illustrative embodiments to be described below, it is assumed without limitation that the initiators of a plurality of initiator-target pairs comprise respective ports of the given host and that the targets of the plurality of initiator-target pairs comprise respective ports of the distributed storage system. Examples of such host ports and storage array ports are illustrated in conjunction with the embodiment of. The host ports can comprise, for example, ports of single-port HBAs and/or ports of multi-port HBAs, or other types of host ports, including network interface cards (NICs). A wide variety of other types and arrangements of initiators and targets can be used in other embodiments.
102 Selecting a particular one of multiple available paths for delivery of a selected one of the IO operations from the given host is more generally referred to herein as “path selection.” Path selection as that term is broadly used herein can in some cases involve both selection of a particular IO operation and selection of one of multiple possible paths for accessing a corresponding logical device of the distributed storage system. The corresponding logical device illustratively comprises a LUN or other logical storage volume to which the particular IO operation is directed.
101 102 100 102 102 106 102 It should be noted that paths may be added or deleted between the hostsand the distributed storage systemin the system. For example, the addition of one or more new paths from the given host to the distributed storage systemor the deletion of one or more existing paths from the given host to the distributed storage systemmay result from respective addition or deletion of at least a portion of the storage devicesof the distributed storage system.
102 Addition or deletion of paths can also occur as a result of zoning and masking changes or other types of storage system reconfigurations performed by a storage administrator or other user. Some embodiments are configured to send a predetermined command from the given host to the distributed storage system, illustratively utilizing the MPIO driver, to determine if zoning and masking information has been changed. The predetermined command can comprise, for example, a log sense command, a mode sense command, a “vendor unique command” or VU command, or combinations of multiple instances of these or other commands, in an otherwise standardized command format.
In some embodiments, paths are added or deleted in conjunction with addition of a new storage array or deletion of an existing storage array from a storage system that includes multiple storage arrays, possibly in conjunction with configuration of the storage system for at least one of a migration operation and a replication operation.
For example, a storage system may include first and second storage arrays, with data being migrated from the first storage array to the second storage array prior to removing the first storage array from the storage system.
As another example, a storage system may include a production storage array and a recovery storage array, with data being replicated from the production storage array to the recovery storage array so as to be available for data recovery in the event of a failure involving the production storage array.
In these and other situations, path discovery scans may be repeated as needed in order to discover the addition of new paths or the deletion of existing paths.
A given path discovery scan can be performed utilizing known functionality of conventional MPIO drivers, such as PowerPath® drivers.
102 102 The path discovery scan in some embodiments may be further configured to identify one or more new LUNs or other logical storage volumes associated with the one or more new paths identified in the path discovery scan. The path discovery scan may comprise, for example, one or more bus scans which are configured to discover the appearance of any new LUNs that have been added to the distributed storage systemas well to discover the disappearance of any existing LUNs that have been deleted from the distributed storage system.
The MPIO driver of the given host in some embodiments comprises a user-space portion and a kernel-space portion. The kernel-space portion of the MPIO driver may be configured to detect one or more path changes of the type mentioned above, and to instruct the user-space portion of the MPIO driver to run a path discovery scan responsive to the detected path changes. Other divisions of functionality between the user-space portion and the kernel-space portion of the MPIO driver are possible. The user-space portion of the MPIO driver is illustratively associated with an Operating System (OS) kernel of the given host.
102 For each of one or more new paths identified in the path discovery scan, the given host may be configured to execute a host registration operation for that path. The host registration operation for a given new path illustratively provides notification to the distributed storage systemthat the given host has discovered the new path.
105 102 101 As indicated previously, the storage nodesof the distributed storage systemprocess IO operations from one or more hostsand in processing those IO operations run various storage application processes that generally involve interaction of that storage node with one or more other ones of the storage nodes.
100 The manner in which functionality for fault domain identifier processing is implemented in systemwill now be described in more detail.
As indicated previously, in distributed storage system arrangements, it can be unduly difficult to discover adequate information regarding targets of respective storage nodes, particularly when using advanced storage access protocols such as NVMe-oF or NVMe/TCP. For example, conventional approaches can lead to sub-optimal arrangements in terms of resiliency and load balancing, thereby adversely impacting storage system performance.
Illustrative embodiments disclosed herein provide techniques for host device processing of fault domain identifiers obtained from respective storage targets in a distributed storage system or other type of storage system. Such techniques can advantageously facilitate host-side functionality such as generating fault domain connectivity information and/or performing path selection, particularly when utilizing advanced storage access protocols such as NVMe-oF and NVMe/TCP, as well as in other storage contexts.
For example, some embodiments utilize the disclosed fault domain identifier processing techniques to provide enhanced resiliency and load balancing in delivery of IO operations from one or more host devices to a storage array or other storage system over paths through a network. These and other embodiments can significantly improve storage system performance.
In some embodiments, an NVMe target, which is an example of what is more generally referred to herein as a “storage target” or simply a “target,” illustratively comprises one or more NVMe controllers, each having a set of TCP connections associated with an administrative (“Admin”) queue and one or more storage-side IO queues. Each TCP connection illustratively corresponds to a single queue having associated request/response entries. In some embodiments, the TCP connections corresponding to a given Admin queue and a set of one or more storage-side IO queues are collectively referred to as a “TCP association.” Typically, a fixed number of storage-side IO queues are established, with a corresponding fixed number of TCP connections, in accordance with maximum load requirements of one or more applications that will be directing IO operations to the NVMe target for processing. The storage-side IO queues and their corresponding TCP connections are examples of what are more generally referred to herein as “network resources.” Such network resources in some embodiments are illustratively used to receive IO operations directed to targets of a storage system from initiators of one or more host devices. Additional or alternative network resources can be used in other embodiments.
The above-noted fault domain identifier processing in some embodiments is illustratively implemented in the following manner.
101 102 104 A given one of the hosts, also referred to herein as a host device, is configured to send commands from the host device to respective ones of a plurality of storage targets of the distributed storage systemover the network. The host device receives in response to the commands, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target. The host device extracts the fault domain identifiers of the respective storage targets from the received storage target data structures, and stores the extracted fault domain identifiers of the respective storage targets in the host device in association with network addresses of the respective storage targets.
The term “data structure” as used herein is intended to be broadly construed, and should not be viewed as being limited to any particular data format. Any of a wide variety of data structures may be used in illustrative embodiments, and a given such data structure may comprise, for example, a combination of multiple smaller data structures, or a portion of a larger data structure.
The network addresses in some embodiments include respective transport addresses of the respective storage targets, where the transport addresses may comprise respective IP addresses obtained from discovery log pages of the storage targets, although other types of network addresses can be used in other embodiments. For example, a given NVMe discovery log page includes a plurality of fields, including bytes 32:63 comprising a Transport Service Identifier (TRSVCID) field that contains a TCP port number for the storage target, and bytes 512:767 comprising a Transport Address (TRADDR) field that provides the IP address for the storage target, in accordance with the NVMe standard.
Storing the extracted fault domain identifiers in association with the network addresses of the respective storage targets is intended to be broadly construed, and may comprise, for example, combining, interconnecting or otherwise associating the extracted fault domain identifiers with respective network addresses in one or more data structures stored in a memory of the host device. As a more particular example, a separate data structure may be generated that includes information for each of at least a subset of a plurality of discovered storage targets, including for each such storage target at least the extracted fault domain identifier and the corresponding network address illustratively obtained from a discovery log page. These and other example arrangements of storing extracted fault domain identifiers in association with their respective network addresses allows the host device to keep track of which connected transport address corresponds to which fault domain.
In some embodiments, the extracted fault domain identifiers are illustratively utilized in the host device to generate fault domain connectivity information and/or to perform path selection. Additional or alternative fault domain processing functionality can be implemented using the extracted fault domain identifiers in other embodiments.
102 The storage targets in some embodiments each comprise one or more NVMe controllers of the distributed storage system, although other types of storage targets can be used in other embodiments.
105 102 102 102 A given fault domain identified by a corresponding one of the extracted fault domain identifiers in some embodiments illustratively comprises at least one storage node of the storage nodesof the distributed storage system, at least one physical server of a plurality of physical servers of the distributed storage system, and/or at least one subnetwork of a plurality of subnetworks of the distributed storage system. Other types of fault domains may be used in other embodiments.
In some embodiments, the commands sent from the host device to respective ones of the storage targets comprise respective NVMe Identify commands, and the storage target data structures comprise respective NVMe Identify Controller data structures, each illustratively modified as disclosed herein to include a fault domain identifier for the corresponding storage target, although other types of commands and storage target data structures may be used in other embodiments, and terms such as “command” and “data structure” as used herein are intended to be broadly construed.
At least a portion of a given one of the fault domain identifiers in some embodiments is illustratively extracted from at least one reserved field of the corresponding storage target data structure, such as one or more reserved fields of the above-noted NVMe Identify Controller data structure.
Additionally or alternatively, at least a portion of a given one of the fault domain identifiers in some embodiments is illustratively extracted from at least one vendor-specific field of the corresponding storage target data structure, such as one or more vendor-specific fields of the above-noted NVMe Identify Controller data structure.
In some embodiments, the host device is further configured to generate fault domain connectivity information based at least in part on the extracted fault domain identifiers, with the fault domain connectivity information indicating different types of network connectivity between the host device and respective fault domains identified by the extracted fault domain identifiers, and to perform one or more automated actions based at least in part on the generated fault domain connectivity information. The one or more automated actions may comprise, for example, controlling initiation of a replication process by the host device and/or generating an alert specifying an absence of a particular type of network connectivity to at least one of the fault domains.
Additionally or alternatively, the host device in some embodiments is further configured to control path selection for delivery of IO operations from the host device to the storage targets based at least in part on the extracted fault domain identifiers. Controlling path selection can comprise, for example, avoiding selection of multiple paths to respective storage targets that have the same fault domain identifier, preventing selection of multiple paths to respective storage targets that have the same fault domain identifier, and/or load balancing delivery of a plurality of IO operations over respective ones of a plurality of paths to respective ones of multiple storage targets each having a different fault domain identifier in a manner that ensures that each of the plurality of IO operations is delivered over a path to a storage target having a different fault domain identifier.
102 102 In some embodiments, the extracted fault domain identifiers are utilized by the host device to match multiple network addresses for respective different storage targets to a same fault domain of the distributed storage system. Such information allows the host device, for example, to achieve improved resiliency and/or load balancing in delivery of IO operations from the host device to the distributed storage system.
3 FIG. A more particular example of an illustrative fault domain identifier processing arrangement of the type described above will be described below in conjunction with. Other types and arrangements of system components supporting fault domain identifier processing can be used in other embodiments.
Additional aspects of some illustrative embodiments will now be described.
105 102 101 As mentioned above, each of the storage nodesof the distributed storage systemillustratively comprises one or more targets, where each such target is associated with multiple distinct paths from respective HBAs or other initiators of one or more of the hosts.
105 102 105 For example, in some embodiments, one or more of the storage nodeseach implements at least one target, such as an NVMe target, that is configured to include multiple controllers, such as at least a first controller associated with a first storage pool, and a second controller associated with a second storage pool. The first and second storage pools are illustratively storage pools of the distributed storage system, and such storage pools may be distributed across multiple ones of the storage nodes. Each of the first and second storage pools is assumed to comprise one or more LUNs or other logical storage volumes.
Although first and second controllers are referred to in conjunction with some embodiments herein, it is to be appreciated that more than two controllers can be implemented in a given target in order to support more than two storage pools.
105 101 101 101 A given one of the storage nodesillustratively processes IO operations received from one or more of the hosts, with different ones of the IO operations being directed by the one or more hostsfrom one or more initiators of the one or more hoststo different ones of the first and second controllers of the target implemented within the given storage node.
102 In some embodiments, each of the storage-side IO queues or set of storage-side IO queues configured in the distributed storage systemfor the given target is associated with a corresponding different TCP connection between the given host and the given target.
105 102 The given target may comprise at least one NVMe controller of a particular one of the storage nodesof the distributed storage system, although other types of targets can be used.
The term “target” as used herein in the context of a distributed storage system or other type of storage system is intended to be broadly construed.
The target in some embodiments more particularly comprises multiple controllers accessible via respective different associations comprising one or more TCP connections between the given host and the given storage node. For example, the target may comprise a plurality of NVMe controllers of an NVMe subsystem of the given storage node.
105 The other storage nodesare each assumed to be configured in a manner similar to that described above and elsewhere herein for the given storage node.
As indicated above, in some embodiments, multiple controllers are part of a single physical controller subsystem of the given storage node. For example, first and second controllers may comprise respective NVMe controllers of an NVMe subsystem of the given storage node. Such an NVMe subsystem is considered an example of what is more generally referred to herein as a “target” of the given storage node.
The first and second controllers in some embodiments may be viewed as comprising respective “virtual” controllers associated with the single physical controller subsystem of the given storage node.
101 Additionally or alternatively, the first and second controllers in some embodiments are accessible via respective first and second different associations comprising one or more TCP connections between a given one of the one or more hostsand the given storage node. In such an arrangement, a host accesses the first controller using the first association, and accesses the second controller using the second association. Such associations are also referred to herein as TCP associations, and may include, for each of at least one Admin queue and a plurality of storage-side IO queues, a corresponding TCP connection. Other types of communication links can be used in other embodiments.
101 In some embodiments, the first controller comprises a first set of IO queues and the second controller comprises a second set of IO queues, for use in processing IO operations for their respective storage pools. Again, each IO queue in a given such set of IO queues may be associated with a separate TCP connection over which a given one of the hostscommunicates with the corresponding controller.
2 FIG. An example of an illustrative process for implementing at least some of the above-described fault domain identifier processing functionality will be provided below in conjunction with the flow diagram of.
105 As indicated previously, the storage nodescollectively comprise an example of a distributed storage system. The term “distributed storage system” as used herein is intended to be broadly construed, so as to encompass, for example, scale-out storage systems, clustered storage systems or other types of storage systems distributed over multiple storage nodes.
Also, the term “storage volume” as used herein is intended to be broadly construed, and should not be viewed as being limited to any particular format or configuration.
A wide variety of alternative configurations of nodes and processing modules are possible in other embodiments. Also, the term “storage node” as used herein is intended to be broadly construed, and may comprise a node that implements storage control functionality but does not necessarily incorporate storage devices. As mentioned previously, a given storage node can in some embodiments comprise a separate storage array, or a portion of a storage array that includes multiple such storage nodes.
1 FIG. The particular features described above in conjunction withshould not be construed as limiting in any way, and a wide variety of other system arrangements implementing fault domain identifier processing as disclosed herein are possible.
105 102 1 FIG. The storage nodesof the example distributed storage systemillustrated inare assumed to be implemented using at least one processing platform, with each such processing platform comprising one or more processing devices, and each such processing device comprising a processor coupled to a memory. Such processing devices can illustratively include particular arrangements of compute, storage and network resources.
105 101 105 The storage nodesmay be implemented on respective distinct processing platforms, although numerous other arrangements are possible. At least portions of their associated hostsmay be implemented on the same processing platforms as the storage nodesor on separate processing platforms.
100 100 101 105 105 101 The term “processing platform” as used herein is intended to be broadly construed so as to encompass, by way of illustration and without limitation, multiple sets of processing devices and associated storage systems that are configured to communicate over one or more networks. For example, distributed implementations of the systemare possible, in which certain components of the system reside in one data center in a first geographic location while other components of the system reside in one or more other data centers in one or more other geographic locations that are potentially remote from the first geographic location. Thus, it is possible in some implementations of the systemfor different subsets of the hostsand the storage nodesto reside in different data centers. Numerous other distributed implementations of the storage nodesand their respective associated sets of hostsare possible.
6 7 FIGS.and Additional examples of processing platforms utilized to implement storage systems and possibly their associated hosts in illustrative embodiments will be described in more detail below in conjunction with.
It is to be appreciated that these and other features of illustrative embodiments are presented by way of example only, and should not be construed as limiting in any way.
101 102 105 106 108 110 112 114 110 105 Accordingly, different numbers, types and arrangements of system components such as hosts, distributed storage system, storage nodes, storage devices, storage processors, system management nodes, control plane serversand instances of path selection logiccan be used in other embodiments. For example, as mentioned previously, system management functionality of system management nodescan be distributed across a subset of the storage nodes, instead of being implemented on separate nodes.
1 FIG. It should therefore be understood that the particular sets of modules and other components implemented in a distributed storage system as illustrated inare presented by way of example only. In other embodiments, only subsets of these components, or additional or alternative sets of components, may be used, and such components may exhibit alternative functionality and configurations.
For example, in other embodiments, certain portions of fault domain identifier processing functionality as disclosed herein can be implemented in one or more hosts, in a storage system, or partially in a host and partially in a storage system. Accordingly, illustrative embodiments are not limited to arrangements in which fault domain identifier processing functionality is implemented primarily in storage system or primarily in a particular host or set of hosts, and therefore such embodiments encompass various alternative arrangements, such as, for example, an arrangement in which the functionality is distributed over one or more storage systems and one or more associated hosts, each comprising one or more processing devices. The term “at least one processing device” as used herein is therefore intended to be broadly construed.
100 102 101 2 FIG. The operation of the information processing systemwill now be described in further detail with reference to the flow diagram of the illustrative embodiment of, which illustrates a process for fault domain identifier processing as disclosed herein. This process may be viewed as an example algorithm implemented at least in part by distributed storage systeminteracting with one or more of the hosts. These and other algorithms for fault domain identifier processing as disclosed herein can be implemented using other types and arrangements of system components in other embodiments.
2 FIG. 200 206 101 105 102 101 105 The process illustrated inincludes stepsthrough, and in some implementations may be performed primarily by at least a given one of the hostsinteracting with at least a subset of the storage nodesof the distributed storage system. Similar processes may be performed primarily by other ones of the hostsinteracting with at least a subset of the storage nodes, although it is to be appreciated that other implementations are possible in other embodiments.
200 In step, commands are sent from a host device to respective ones of a plurality of storage targets of a storage system over a network. In some embodiments, the commands sent from the host device to respective ones of the storage targets comprise respective NVMe Identify commands of an NVMe access protocol, including NVMe-oF or NVMe/TCP, although the disclosed techniques are applicable for use with other storage access protocols, including SCSI and iSCSI access protocols.
Prior to sending the commands, the host device in some embodiments discovers the storage targets using respective discovery log pages, illustratively obtained from one or more discovery controllers which in some embodiments are implemented on one or more of the storage targets and/or on other storage nodes or processing devices of the storage system, although other arrangements are possible. In some embodiments, such discovery log pages are illustratively obtained using corresponding commands of a storage access protocol, such as the NVMe access protocol. However, the term “discovery log page” as used herein is intended to be broadly construed, as illustratively comprising, for example, at least one log page or a suitable portion thereof that includes one or more entries comprising information utilized by the host to discover and connect to at least one corresponding storage target, and should not be viewed as being limited to a particular type of discovery log page configured in accordance with a particular storage access protocol.
The storage targets in some embodiments each comprise one or more NVMe controllers of the storage system, although other types of storage targets can be used in other embodiments. As indicated elsewhere herein, terms such as “storage target” and “target” as used herein are intended to be broadly construed, and in some embodiments can comprise, for example, an NVMe subsystem, which is more generally referred to herein as an NVMe target. The NVMe subsystem or other NVMe target in such an arrangement illustratively comprises one or more controllers.
202 In step, in response to the commands, the host device receives, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target.
In some embodiments, the storage target data structures comprise respective NVMe Identify Controller data structures, each illustratively modified as disclosed herein to include a fault domain identifier for the corresponding storage target, although other types of storage target data structures may be used in other embodiments.
The fault domain identifiers indicate respective fault domains of the storage system. A given one of the fault domains in some embodiments illustratively comprises at least one storage node of a plurality of storage nodes of the storage system, at least one physical server of a plurality of physical servers of the storage system, and/or at least one subnetwork of a plurality of subnetworks of the storage system.
Each of the fault domain identifiers therefore illustratively indicates a fault domain of the corresponding storage target. For example, the fault domain may comprise a particular storage node of a plurality of storage nodes of the storage system, and the fault domain identifier may comprise an identifier of the particular storage node. As another example, the fault domain may comprise a particular physical server of a plurality of physical servers of the storage system, and the fault domain identifier may comprise an identifier of the particular physical server. As yet another example, the fault domain may comprise a particular subnetwork of a plurality of subnetworks of the storage system, and the fault domain identifier may comprise an identifier of the particular subnetwork.
These are only examples, and a wide variety of other arrangements of fault domains and associated fault domain identifiers may be used in other embodiments. In some embodiments, the fault domains comprise respective failure domains of the storage system. The term “fault domain” as used herein is therefore intended to be broadly construed, so as to encompass, for example, failure domains, storage-side application domains, and numerous other types of fault domains of the storage system.
204 In step, the fault domain identifiers of the respective storage targets are extracted from the received storage target data structures by the host device.
4 FIG. In some embodiments, at least a portion of a given one of the fault domain identifiers in some embodiments is illustratively extracted from at least one reserved field of the corresponding storage target data structure, such as one or more reserved fields of the above-noted NVMe Identify Controller data structure. A more particular example of an arrangement of this type will be described in conjunction withbelow.
Additionally or alternatively, at least a portion of a given one of the fault domain identifiers in some embodiments is illustratively extracted from at least one vendor-specific field of the corresponding storage target data structure, such as one or more vendor-specific fields of the above-noted NVMe Identify Controller data structure.
It is to be appreciated, however, that a wide variety of other types of fault domain identifiers and associated storage target data structures may be used, and terms such as “fault domain identifier” and “storage target data structure” as used herein are therefore intended to be broadly construed.
206 In step, the extracted fault domain identifiers of the respective storage targets are stored in the host device in association with network addresses of the respective storage targets.
The network addresses in some embodiments include respective transport addresses of the respective storage targets, where the transport addresses may comprise respective IP addresses obtained from the above-noted discovery log pages of the storage targets. Other types of network addresses can be used in other embodiments.
As indicated above, in some embodiments, the extracted fault domain identifiers are illustratively utilized in the host device to generate fault domain connectivity information and/or to perform path selection. Additional or alternative fault domain processing functionality can be implemented using the extracted fault domain identifiers in other embodiments.
For example, in some embodiments, the host device generates fault domain connectivity information based at least in part on the extracted fault domain identifiers, with the fault domain connectivity information indicating different types of network connectivity between the host device and respective fault domains, and performs one or more automated actions based at least in part on the generated fault domain connectivity information.
The one or more automated actions may comprise, for example, controlling initiation of a replication process by the host device and/or generating an alert specifying an absence of a particular type of network connectivity to at least one of the fault domains.
Additionally or alternatively, the host device in some embodiments controls path selection for delivery of IO operations from the host device to the storage targets based at least in part on the extracted fault domain identifiers.
Controlling path selection can comprise, for example, avoiding selection of multiple paths to respective storage targets that have the same fault domain identifier, preventing selection of multiple paths to respective storage targets that have the same fault domain identifier, and/or load balancing delivery of a plurality of IO operations over respective ones of a plurality of paths to respective ones of multiple storage targets each having a different fault domain identifier in a manner that ensures that each of the plurality of IO operations is delivered over a path to a storage target having a different fault domain identifier.
Such path selection control functionality is illustratively implemented at least in part in path selection logic of one or more corresponding MPIO drivers of the host. Other types of path selection control can be implemented based at least in part on extracted fault domain identifiers as disclosed herein.
In these and other arrangements, the extraction of the fault domain identifiers from the storage target data structures advantageously allows the host to match multiple network addresses for respective different storage targets to a same fault domain of the storage system. Such information is utilized by the host to improve resiliency and/or load balancing in the delivery of IO operations to the targets, as previously described.
200 206 One or more of stepsthroughare illustratively repeated over time in order to support the fault domain identifier processing functionality disclosed herein. Multiple such processes may operate in parallel with one another in order to provide fault domain identifier processing functionality for different host devices and/or different sets of storage targets of the storage system.
2 FIG. The steps of theprocess are shown in sequential order for clarity and simplicity of illustration only, and certain steps can at least partially overlap with other steps. Additional or alternative steps can be used in other embodiments.
2 FIG. The particular processing operations and other system functionality described in conjunction with the flow diagram ofare therefore presented by way of illustrative example only, and should not be construed as limiting the scope of the disclosure in any way. Alternative embodiments can use other types of processing operations for implementing fault domain identifier processing for one or more hosts and a storage system. For example, as indicated above, the ordering of the process steps may be varied in other embodiments, or certain steps may be performed at least in part concurrently with one another rather than serially. Also, one or more of the process steps may be repeated periodically, or multiple instances of the process can be performed in parallel with one another in order to implement a plurality of different processes for respective different hosts and/or sets of storage targets.
2 FIG. Functionality such as that described in conjunction with the flow diagram ofcan be implemented at least in part in the form of one or more software programs stored in memory and executed by a processor of a processing device such as a computer or server. As will be described below, a memory or other storage device having executable program code of one or more software programs embodied therein is an example of what is more generally referred to herein as a “processor-readable storage medium.”
One or more hosts and/or one or more storage nodes can be implemented as part of what is more generally referred to herein as a processing platform comprising one or more processing devices each comprising a processor coupled to a memory.
A given such processing device in some embodiments may correspond to one or more virtual machines or other types of virtualization infrastructure such as Docker containers or Linux containers (LXCs). Hosts, storage processors and other system components may be implemented at least in part using processing devices of such processing platforms. For example, respective path selection logic instances and other related logic instances of the hosts can be implemented in respective containers running on respective ones of the processing devices of a processing platform.
3 5 FIGS.through Additional examples of illustrative embodiments will now be described with reference to.
In some embodiments, a storage system based on the NVMe access protocol is configured as a distributed system with multiple NVMe targets, each providing a different storage system interface. Application servers or other host devices can connect to several of the NVMe targets for IO load balancing, bandwidth utilization and congestion avoidance, illustratively utilizing path selection algorithms implemented in MPIO drivers of the type noted above.
An important issue that arises in the context of path selection in an MPIO driver relates to how to distribute IO load between various paths to different ones of the NVMe targets, and how to avoid multiple path failures in situations in which multiple paths connect to NVMe targets in the same fault domain.
For example, it is possible that two separate paths of a given multi-path arrangement for reaching distinct NVMe targets, as reported in one or more discovery log pages sent from the NVMe targets to an NVMe initiator of a given host device, are going to the same physical server or other storage node. For both resiliency and load balancing, the MPIO driver of the given host device should select and connect only to one of the NVMe targets on the same physical server. However, the current NVMe standard does not provide sufficient information to allow the MPIO driver or other host component to make this determination.
Accordingly, illustrative embodiments disclosed herein extend a storage target data structure, such as an NVMe Identify Controller data structure, to include a fault domain identifier that identifies the fault domain of the corresponding storage target. The particular storage implementation can determine what storage system entities the fault domain identifier represents. For example, it can be used to represent a particular fault domain, such as a storage node, a physical server, a subnetwork, an equipment rack, or other storage system configuration entity, as determined by a storage system administrator or other user. As a more particular example, the fault domain identifier can comprise a server tag that is unique for each physical server of the storage system, and included in a storage target data structure returned from an NVMe target of the storage system to an NVMe initiator on the host device.
More generally, the fault domain identifier in some embodiments is configured to represent a general fault domain, against which the given host device will provide various types of fault domain identifier processing functionality, such as generating connectivity information and/or performing path selection. These and other types of functionality implemented using the techniques disclosed herein can significantly improve resiliency and load balancing in the storage system.
In some embodiments, a path selection algorithm implemented by the MPIO driver of the given host device is illustratively configured to load balance the delivery of IO operations over paths to the storage system across different storage system fault domains, such as storage nodes, physical servers, subnetworks or other storage system entities represented by the fault domain identifier as disclosed herein.
Such an arrangement advantageously allows the MPIO driver to consider the fault domain identifiers when determining which storage system controller or other NVMe target to connect to for delivery of a given IO operation. For example, the MPIO driver can be configured to avoid connecting to two NVMe targets in the same fault domain. This can ensure maximal resilience when the host device is connected to NVMe targets represented by separate fault domains, by avoiding potential double failures that might otherwise result in cases in which multiple paths are directed to the same fault domain of the storage system. Moreover, optimal load balancing is provided when the paths of the multi-path arrangement go to respective different physical servers or other respective different fault domains of the storage system.
Illustrative embodiments allow the host device and its MPIO driver to better configure the paths to NVMe targets of the storage system, by allowing the host device and its MPIO driver to determine in advance, from the above-described fault domain identifiers inserted in storage target data structures, that certain NVMe targets are part of the same fault domain, such as storage node, physical server, subnetwork or other storage system entity relevant to resilience and/or load balancing. Moreover, the need for creation of redundant connections, and their associated time-consuming connect, identify controller and disconnect operations, is advantageously avoided in illustrative embodiments. For example, using the disclosed techniques, there is no need to connect to and issue identify controller commands to an NVMe target just to learn that the NVMe target is part of the same physical server.
3 FIG. Referring now to, this embodiment illustrates an example of a distributed storage system that more particularly comprises a software-defined storage system having a plurality of software-defined storage server nodes, also referred to as SDS server nodes, configured to utilize an NVMe storage access protocol such as NVMe-oF or NVMe/TCP. Such SDS server nodes are examples of “storage nodes” as that term is broadly used herein. As will be appreciated by those skilled in the art, similar embodiments can be implemented without the use of software-defined storage and with other storage access protocols.
3 FIG. 300 301 302 305 1 305 2 305 3 305 4 305 As shown in, an information processing systemcomprises a hostconfigured to communicate over a network, not explicitly shown but illustratively a TCP/IP network, with a software-defined storage system comprising an NVMe target clustercomprising four distinct storage servers-,-,-and-. The storage serversare illustratively implemented as respective SDS server nodes, although numerous other types of storage servers can be used.
311 301 305 318 301 314 315 301 315 314 314 315 301 301 300 300 301 100 1 FIG. A plurality of applicationsexecute on the hostand generate IO operations that are delivered to particular ones of the storage serversvia at least one NVMe initiator. The hostfurther comprises path selection logicand fault domain identifier processing logic, illustratively configured to carry out aspects of fault domain identifier processing functionality of the hostin a manner similar to that previously described. In other embodiments, the fault domain identifier processing logicmay be part of the path selection logic, rather than a separate component as illustrated in the figure. Both the path selection logicand the fault domain identifier processing logicin some embodiments are implemented at least in part within an MPIO driver of the host. Although only a single hostis shown in system, the systemcan include multiple hosts, each configured as generally shown for host, as in the systemof.
305 1 FIG. Each of the storage serversin the present embodiment comprises at least one NVMe target, and may include additional components, such as a data relay agent, a data server and a set of local drives. The internal components of a given storage server with the exception of the local drives are illustratively part of a corresponding storage processor in theembodiment, although numerous other arrangements are possible.
305 305 305 The data relay agent facilitates relaying of IO requests between different ones of the storage servers, and the data servers provide access to data stored in the local drives of their respective storage servers. Additional or alternative components may be included in the storage serversin illustrative embodiments.
3 FIG. 318 301 301 301 305 In theembodiment, the NVMe initiatorof the hostillustratively sends commands to respective ones of a plurality of discovered storage targets, such as NVMe Identify commands, and receives in response to those commands, from each of the storage targets, a corresponding storage target data structure, illustratively an NVMe Identify Controller data structure, that includes in one or more fields thereof a fault domain identifier as disclosed herein, which illustratively identifies a corresponding fault domain of the storage system. Although a single NVMe initiator is shown in the host, this is by way of simplified illustration only, and other embodiments can include multiple NVMe initiators within host. Each of the storage serverscan similarly include multiple NVMe targets. Also, one or more instances of fault domain identifier processing logic are assumed to be implemented within or otherwise associated with each of the NVMe targets, although such storage-side logic instances are not shown in the figure.
305 In some embodiments, the storage serversare configured at least in part as respective PowerFlex® software-defined storage nodes from Dell Technologies, suitably modified as disclosed herein to implement fault domain identifier processing, although other types of storage nodes can be used in other embodiments.
The NVMe targets in some embodiments collectively comprise an NVMe subsystem that implements multiple distinct controllers and associated aspects of the fault domain identifier processing. For example, a given such NVMe target can comprise at least a first controller associated with a first storage pool of the distributed storage system, and a second controller associated with a second storage pool of the distributed storage system. Other types and arrangements of multiple controllers can be used.
305 301 301 314 315 318 301 305 302 A given one of the storage serversprocesses IO operations received from the host, with different ones of the IO operations being directed by the host, at least in part under the control of path selection logicand fault domain identifier processing logic, from the NVMe initiatorof the hostto different ones of the NVMe targets of the storage serversin the NVMe target cluster.
302 305 1 305 2 305 3 305 4 1 2 3 4 305 1 2 1 3 2 4 3 1 2 3 The storage system in this example is assumed to include NVMe target clustercomprising the four physical servers-,-,-and-, also respectively denoted herein as Storage Server, Storage Server, Storage Server, and Storage Server. These storage serversare distributed across three different fault domains as shown in the figure, with Storage Serverand Storage Serverbeing in Fault Domain, Storage Serverbeing in Fault Domain, and Storage Serverbeing in Fault Domain. Fault Domainmay be viewed as an example of a fault domain that comprises multiple physical servers, possibly collectively comprising a subnetwork, while Fault Domainsandare examples of fault domains that each comprise a single physical server.
318 301 302 305 1 2 3 4 1 2 3 4 1 1 2 2 3 3 4 4 The NVMe initiatorof the hostreceives four discovery log page entries respectively indicating four different multi-path interfaces for four storage servers of the NVMe target cluster, including one interface for each of the storage servers. The interfaces are associated with multiple paths denoted mpath-, mpath-, mpath-and mpath-as shown, also referred to as simply path, path, pathand path, respectively. In this example multi-path arrangement, pathgoes to Storage Server, pathgoes to Storage Server, pathgoes to Storage Serverand pathgoes to Storage Server.
318 301 301 301 The NVMe initiatorof the hostillustratively connects to a discovery controller on an NVMe target to obtain at least a portion of the discovery log pages. Additional or alternative techniques can be used to obtain discovery log pages in hostin other embodiments. For example, one or more discovery controllers associated with a discovery service implemented at least in part on one or more NVMe subsystems or other NVMe targets can provide one or more of the discovery log pages to the hostin illustrative embodiments.
301 It is important in a variety of different processing contexts for the hostto be aware of the fault domains in which the various storage interfaces reside. For example, in the context of a replication process, in which data of one or more storage volumes is replicated from one storage cluster to another storage cluster, there may be restrictions in terms of what fault domains are appropriate replication destinations for the one or more storage volumes. Such a restriction may provide, for example, that certain storage volumes are to be replicated only to particular fault domains of the storage system.
302 Absent use of the fault domain identifier processing techniques disclosed herein, a host device and its MPIO driver might otherwise be unaware that multiple reported interfaces of the NVMe target clusterare in the same fault domain, leading to potential problems in replicating data of one or more storage volumes as well as in numerous other contexts relating to resilience and/or load balancing.
301 302 301 Illustrative embodiments overcome these and other drawbacks by, for example, allowing the hostand its MPIO driver to accurately and efficiently determine in advance which storage targets are associated with which fault domains of the NVMe target cluster. The hostcan then utilize the fault domain identifiers to provide processing functionality such as generating fault domain connectivity information and/or controlling path selection.
In some embodiments, the fault domain identifiers disclosed herein are implemented through a change in the NVMe standard, with the addition of a fault domain identifier field, for example, in one or more reserved fields and/or vendor-specific fields of an NVMe Identify Controller data structure. Additional details regarding NVMe Identify commands, NVMe Identify Controller data structures and other aspects of the NVMe standard can be found in, for example, the above-cited NVM Express Base Specification, Revision 2.1, August 2024, and its associated NVM Express Command Set Specification and NVM Express TCP Transport Specification, although other NVMe implementations can be used.
4 FIG. 400 400 401 402 400 shows an example of a storage tag data structure in an illustrative embodiment. This example storage tag data structure more particularly comprises an NVMe Identify Controller data structuregenerally configured accordance with the NVMe standard, but modified as disclosed herein to include a 64-bit fault domain identifier in at least one reserved field comprising the eight bytes 388:395 of the NVMe Identify Controller data structure. More particularly, a first portionof the 64-bit fault domain identifier (“ID”) occupies the 32 bits of the four bytes 388:391 and a second portionof the 64-bit fault domain identifier occupies the 32 bits of the four bytes 392:395. In accordance with the fault domain identifier processing techniques disclosed herein, a particular storage target, such as an NVMe target comprising a controller, inserts its corresponding 64-bit fault domain identifier into the eight bytes 388:395, before sending the NVMe Identify Controller data structureto the host device in response to an NVMe Identify command received from that host device.
The 64-bit fault domain identifier carried by the eight bytes 388:395 illustratively indicates the particular type of fault domain, such as whether the fault domain is a particular storage node or set of storage nodes, a physical server or set of physical servers, or a particular subnetwork or set of subnetworks, as well as further identifying information such as an identifier of the particular storage node(s), physical server(s) or subnetwork(s). Any of a wide variety of different identifying information formats may be used for the 64-bit fault domain identifier.
400 4 FIG. Other portions of the NVMe Identify Controller data structureare configured in accordance with the existing features of the NVMe standard. For example, other fields include a 16-bit Subsystem Vendor ID (SSVID) field, a 16-bit PCI Vendor ID (PVID) field, a 16-bit Command Quiesce Time (CQT) field, and an 8-bit Temperature Threshold Hysteresis Attributes (TMPTHHA) field. Additional fields of the example storage target data structure ofare similarly as defined in the NVMe standard.
301 315 1 2 1 The above-noted bytes 388:395 are currently reserved under the NVMe standard, but through a change in the standard could be configured to carry a 64-bit fault domain identifier of the type disclosed herein. Such fault domain identifiers would allow the hostand its fault domain identifier processing logicto determine that, in the context of the above example, that Storage Serverand Storage Serviceare part of the same fault domain, namely, Fault Domain, and to perform associated processing such as generating fault domain connectivity information and/or controlling path selection.
For example, some embodiments generate fault domain connectivity information based at least in part on the extracted fault domain identifiers, with the fault domain connectivity information indicating different types of network connectivity between the host device and respective fault domains, and perform one or more automated actions based at least in part on the generated fault domain connectivity information. The different types of network connectivity may comprise at least full connectivity, partial connectivity and no connectivity, although additional or alternative connectivity types may be used.
The automated actions may comprise controlling initiation of a replication process by the host device, such that replication of particular storage volumes is only carried out to particular fault domains in accordance with the above-noted restrictions.
Additionally or alternatively, the automated actions may comprise generating an alert specifying an absence of a particular type of network connectivity to at least one of the fault domains. Such an alert may be sent to a storage administrator or other user so that the storage administrator or other user can address the issue leading to the absence of the particular type of network connectivity. Additionally or alternatively, the alert may be sent to an automated system entity that addresses the absence of the particular type of network connectivity.
In some embodiments, the host device utilizes the extracted fault domain identifiers and their associated network addresses to generate network connectivity information in the form of a detailed connectivity report that characterizes multiple NVMe controllers based at least in part on the fault domains to which those controllers belong. Such a report illustratively includes a current state of network connectivity to particular fault domains (e.g., physical servers or subnetworks comprising multiple servers), including the particular network addresses associated with the network connections of the host device to the particular fault domains. The report can additionally or alternatively include an indication of one or more network addresses for which connections to the corresponding fault domains have not yet been established by the host and/or one or more network addresses for which previous connections have been disconnected.
The network connectivity information generated by the host device in some embodiments therefore illustratively indicates which connections to particular fault domains are currently active and what fault domains are fully connected, partially connected or not connected.
The above-described example connectivity report can be provided in some embodiments as part of an alert sent to a storage administrator or other user. The connectivity report in some embodiments provides an indication to the storage administrator of potential problems in the distributed storage system, so that the storage administrator can address any such issues. For example, the storage administrator can take steps to recover full connectivity for those storage nodes, physical servers and/or subnetworks for which the connectivity report currently indicates the corresponding fault domain connectivity as disconnected, partially connected or never connected.
In some embodiments, the connectivity report is used to enforce replication policies, for example, where certain storage volumes locally connected to the host device must be replicated to a certain fault domain of the distributed storage system. The connectivity report can additionally or alternatively indicate if particular fault domains have full connectivity to the host device with a sufficient number of redundant paths, or instead have only a single connection or no connection at all. The above-noted replication policy enforcement can be done automatically and/or via storage administrator intervention.
It is to be appreciated that a wide variety of other techniques can be utilized to incorporate a fault domain identifier of the type disclosed herein into a storage target data structure with or without modification of the existing NVMe standard.
400 400 4 FIG. As another example, in some embodiments, the fault domain identifier is incorporated into one or more vendor-specific fields of the NVMe Identify Controller data structurerather than using the particular reserved fields as illustrated in. In such an arrangement, the 64-bit fault domain identifier is illustratively incorporated into a portion of bytes 4095:3072 of the NVMe Identify Controller data structure, but its particular exact placement within those bytes would generally be up to each vendor. The length and configuration of the fault domain indicator can also be varied depending on factors such as the implementation and the available fields of the storage target data structure.
Also, similar fault domain identifier processing techniques can be implemented in other types of storage systems utilizing other types of fault domain identifiers, in any of a wide variety of different lengths and configurations.
4 FIG. It should also be understood that the particular features and functionality described above are examples only. Accordingly, the particular fault domain identifier configuration as shown inis presented by way of illustrative example only, and should not be viewed as limiting in any way. A wide variety of different alternative arrangements of storage target data structures and associated fault domain identifiers can be used in other embodiments.
5 FIG. 500 511 514 515 521 522 514 515 521 522 500 500 514 515 515 514 Referring now to, another illustrative embodiment is shown. In this embodiment, an information processing systemcomprises host-side elements that include application processes, path selection logicand fault domain identifier (“ID”) processing logic, and storage-side elements that include multiple targetsand fault domain identifier processing logic. The path selection logicis configured to operate in conjunction with fault domain identifier processing logic, multiple targetsand fault domain identifier processing logicto implement functionality for fault domain identifier processing in the system. There may be separate instances of one or more such elements associated with each of a plurality of system components such as hosts and storage arrays of the system. For example, different instances of the path selection logicand fault domain identifier processing logicare illustratively implemented within or otherwise in association with respective ones of a plurality of MPIO drivers of respective hosts. In other embodiments, the fault domain identifier processing logiccan be implemented at least in part within the path selection logic.
500 530 532 534 536 538 540 530 532 534 536 538 540 The systemis configured in accordance with a layered system architecture that illustratively includes a host processor layer, an MPIO layer, a host port layer, a switch fabric layer, a storage array port layerand a storage array processor layer. The host processor layer, the MPIO layerand the host port layerare associated with one or more hosts, the switch fabric layeris associated with one or more SANs or other types of networks, and the storage array port layerand storage array processor layerare associated with one or more storage arrays (“SAs”). A given such storage array illustratively comprises a software-defined storage system or other type of distributed storage system comprising a plurality of storage nodes.
500 521 In a manner similar to that described elsewhere herein, one or more storage arrays of the systemare each configured to implement one or more storage targets, such as, for example, at least a first controller associated with a first storage pool, and a second controller associated with a second storage pool, where the first and second controllers each include respective sets of IO queues. Numerous other arrangements of multiple targetscan be used.
500 532 514 515 521 522 The systemin this embodiment implements fault domain identifier processing functionality utilizing one or more MPIO drivers of the MPIO layer, and associated instances of path selection logicand fault domain identifier processing logic, as well as the multiple targetsand the fault domain identifier processing logic.
500 521 For example, a given host of the systemis illustratively configured to send commands from the host to respective ones of the multiple targetsover a network, to receive in response to the commands, from each of the storage targets, a corresponding storage target data structure that comprises in one or more fields thereof a fault domain identifier of that storage target, to extract the fault domain identifiers of the respective storage targets from the received storage target data structures, and to store the extracted fault domain identifiers of the respective storage targets in the host device in association with network addresses of the respective storage targets.
The given host will then provide various types of fault domain identifier processing functionality, such as generating connectivity information and/or performing path selection, as described in more detail elsewhere herein. For example, connectivity reports of the type described above may be generated and transmitted with corresponding alerts.
511 514 532 With regard to path selection, for each of a plurality of IO operations generated by one or more of the application processesin the given host for delivery to the given storage array, the given host selects, illustratively via path selection logicof one or more MPIO drivers of the MPIO layer, a particular one of the plurality of paths from the initiator to one of the targets on the particular storage node, and sends the IO operation to the particular storage node over the selected path. Such path selection is based at least in part on the above-noted fault domain identifiers extracted from respective storage target data structures by the given host, as disclosed herein.
511 532 514 The application processesgenerate IO operations that are processed by the MPIO layerfor delivery to the one or more storage arrays that collectively comprise a plurality of storage nodes of a distributed storage system. Paths are determined by the path selection logicfor sending such IO operations to the one or more storage arrays. These IO operations are sent to the one or more storage arrays in accordance with one or more scheduling algorithms, load balancing algorithms and/or other types of algorithms, based at least in part on storage system fault domains as reflected in extracted fault domain identifiers.
532 514 515 The MPIO layeris an example of what is also referred to herein as a multi-path layer, and comprises one or more MPIO drivers implemented in respective hosts. Each such MPIO driver illustratively comprises respective instances of path selection logicand fault domain identifier processing logicconfigured as previously described. Additional or alternative layers and logic arrangements can be used in other embodiments.
514 The one or more storage arrays process IO operations received from one or more hosts, with different ones of the IO operations being directed by the one or more hosts under the control of path selection logicfrom one or more initiators of the one or more hosts to different targets based at least in part on extracted fault domain identifiers as disclosed herein.
521 540 521 522 515 514 The multiple targetsare implemented in the storage array processor layer, with particular logical storage volumes being accessible via respective ones of the targets. Fault domains of respective ones of the multiple targetsare characterized by fault domain identifiers that are inserted into storage target data structures, such as NVMe Identify Controller data structures, illustratively at least in part by fault domain identifier processing logic, and subsequently extracted by fault domain identifier processing logicand utilized to modify path selection in path selection logicand to provide other functionality such as generating connectivity information.
500 514 1 1 1 2 2 2 As mentioned above, in the system, path selection logicis configured to select different paths for sending IO operations from a given host to a storage array. These paths as illustrated in the figure include a first path from a particular host port denoted HPthrough a particular switch fabric denoted SFto a particular storage array port denoted SP, and a second path from another particular host port denoted HPthrough another particular switch fabric denoted SFto another particular storage array port denoted SP.
5 FIG. These two particular paths are shown by way of illustrative example only, and in many practical implementations there will typically be a much larger number of paths between the one or more hosts and the one or more storage arrays, depending upon the specific system configuration and its deployed numbers of host ports, switch fabrics and storage array ports. For example, each host in theembodiment can illustratively have the same number and type of paths to a shared storage array, or alternatively different ones of the hosts can have different numbers and types of paths to the storage array.
514 532 538 514 521 522 515 The path selection logicof the MPIO layerin this embodiment selects paths for delivery of IO operations to the one or more storage arrays having the storage array ports of the storage array port layer. More particularly, the path selection logicdetermines appropriate paths over which to send particular IO operations to particular logical storage devices of the one or more storage arrays. As disclosed herein, such path selection is based at least in part on respective fault domains of the multiple targetsas indicated by fault domain identifiers, which are illustratively inserted by fault domain identifier processing logicand extracted by fault domain identifier processing logic.
500 Some implementations of the systemcan include a relatively large number of hosts (e.g., 1000 or more hosts), although as indicated previously different numbers of hosts, and possibly only a single host, may be present in other embodiments. Each of the hosts is typically allocated a sufficient number of host ports to accommodate predicted performance needs. In some cases, the number of ports per host is on the order of 4, 8 or 16, although other numbers of ports could be allocated to each host depending upon the predicted performance needs. A typical storage array may include on the order of 128 ports, although again other numbers can be used based on the particular needs of the implementation. The number of hosts per storage array port in some cases can be on the order of 10 hosts per port.
500 A given host of systemcan be configured to initiate an automated path discovery process to discover new paths responsive to updated zoning and masking or other types of storage system reconfigurations performed by a storage administrator or other user. For certain types of hosts, such as hosts using particular operating systems such as Windows, ESX or Linux, automated path discovery via the MPIO drivers of a multi-path layer is typically supported. Other types of hosts using other operating systems such as AIX in some implementations do not necessarily support such automated path discovery, in which case alternative techniques can be used to discover paths.
These and other features of illustrative embodiments disclosed herein are examples only, and should not be construed as limiting in any way. Other types of fault domain identifier processing can be used in other embodiments, and the term “fault domain identifier processing” as used herein is intended to be broadly construed.
The above-described illustrative embodiments can provide significant advantages over conventional approaches.
For example, some embodiments provide techniques for fault domain identifier processing, in a software-defined storage system or other type of distributed storage system, to allow one or more host devices to obtain additional information for respective NVMe targets or other types of targets of a storage system. Such information is used by a host device in enhancing resiliency and/or load balancing for delivery of IO operations to the storage system.
Accordingly, some embodiments provide improved resiliency in the presence of failures as well as improved load balancing.
The above-described fault domain identifier functionality advantageously allows the host device to match multiple network addresses for respective different storage targets to a same fault domain of the storage system and to generate network connectivity information and/or adjust its path selection accordingly.
For example, the path selection may be configured such that the host device does not utilize more than one path to the same fault domain entity, especially in high latency networks, where connecting and disconnecting NVMe controllers might otherwise take significant amounts of time. The amount of volume mappings reported over an NVMe controller also contributes to this connect/disconnect time.
The disclosed techniques also conserve system bandwidth, as well as reducing setup time for storage administrators or other users, especially in high latency networks.
By allowing host NVMe initiators to more accurately and efficiently learn which NVMe targets are part of which fault domains, such as storage nodes, physical servers or other storage system entities, optimal path selection configurations can be readily determined and implemented in a host device.
The multi-pathing portions of the example techniques described above may be performed by a given MPIO driver on a corresponding host device, and similarly by other MPIO drivers on respective other host devices. Such MPIO drivers illustratively form a multi-path layer comprising multi-pathing software of the host devices. Other types of host drivers can be used in other embodiments.
Although various types of commands and log pages are used in illustrative embodiments herein, other types of commands and log pages can be used in other embodiments. For example, various types of log sense, mode sense and/or other “read-like” commands, possibly including one or more commands of a standard storage access protocol such as the above-noted NVMe and SCSI access protocols, can be used in other embodiments.
Although some embodiments implement fault domain identifiers through modification of an NVMe specification or other storage access protocol specification, other embodiments can be implemented without requiring any change in the NVMe specification or other storage access protocol specification.
It is to be appreciated that the particular advantages described above and elsewhere herein are associated with particular illustrative embodiments and need not be present in other embodiments. Also, the particular types of information processing system features and functionality as illustrated in the drawings and described above are exemplary only, and numerous other arrangements may be used in other embodiments.
6 7 FIGS.and 100 Illustrative embodiments of processing platforms utilized to implement hosts and distributed storage systems with fault domain identifier processing functionality will now be described in greater detail with reference to. Although described in the context of system, these platforms may also be used to implement at least portions of other information processing systems in other embodiments.
6 FIG. 600 600 100 600 602 1 602 2 602 604 604 605 shows an example processing platform comprising cloud infrastructure. The cloud infrastructurecomprises a combination of physical and virtual processing resources that may be utilized to implement at least a portion of the information processing system. The cloud infrastructurecomprises multiple virtual machines (VMs) and/or container sets-,-, . . .-L implemented using virtualization infrastructure. The virtualization infrastructureruns on physical infrastructure, and illustratively comprises one or more hypervisors and/or operating system level virtualization infrastructure. The operating system level virtualization infrastructure illustratively comprises kernel control groups of a Linux operating system or other type of operating system.
600 610 1 610 2 610 602 1 602 2 602 604 602 The cloud infrastructurefurther comprises sets of applications-,-, . . .-L running on respective ones of the VMs/container sets-,-, . . .-L under the control of the virtualization infrastructure. The VMs/container setsmay comprise respective VMs, respective sets of one or more containers, or respective sets of one or more containers running in VMs.
6 FIG. 602 604 100 In some implementations of theembodiment, the VMs/container setscomprise respective VMs implemented using virtualization infrastructurethat comprises at least one hypervisor. Such implementations can provide fault domain identifier processing functionality in a distributed storage system of the type described above using one or more processes running on a given one of the VMs. For example, each of the VMs can include logic instances and/or other components for implementing functionality associated with fault domain identifier processing in the system.
604 A hypervisor platform may be used to implement a hypervisor within the virtualization infrastructure. Such a hypervisor platform may comprise an associated virtual infrastructure management system. The underlying physical machines may comprise one or more distributed processing platforms that include one or more storage systems.
6 FIG. 602 604 100 In other implementations of theembodiment, the VMs/container setscomprise respective containers implemented using virtualization infrastructurethat provides operating system level virtualization functionality, such as support for Docker containers running on bare metal hosts, or Docker containers running on VMs. The containers are illustratively implemented using respective kernel control groups of the operating system. Such implementations can also provide fault domain identifier processing functionality in a distributed storage system of the type described above. For example, a container host supporting multiple containers of one or more container sets can include logic instances and/or other components for implementing functionality associated with fault domain identifier processing in the system.
100 600 700 6 FIG. 7 FIG. As is apparent from the above, one or more of the processing devices or other components of systemmay each run on a computer, server, storage device or other processing platform element. A given such element may be viewed as an example of what is more generally referred to herein as a “processing device.” The cloud infrastructureshown inmay represent at least a portion of one processing platform. Another example of such a processing platform is processing platformshown in.
700 100 702 1 702 2 702 3 702 704 The processing platformin this embodiment comprises a portion of systemand includes a plurality of processing devices, denoted-,-,-, . . .-K, which communicate with one another over a network.
704 The networkmay comprise any type of network, including by way of example a global computer network such as the Internet, a WAN, a LAN, a satellite network, a telephone or cable network, a cellular network, a wireless network such as a WiFi or WiMAX network, or various portions or combinations of these and other types of networks.
702 1 700 710 712 The processing device-in the processing platformcomprises a processorcoupled to a memory.
710 The processormay comprise a microprocessor, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), graphics processing unit (GPU) or other type of processing circuitry, as well as portions or combinations of such circuitry elements.
712 712 The memorymay comprise random access memory (RAM), read-only memory (ROM), flash memory or other types of memory, in any combination. The memoryand other memories disclosed herein should be viewed as illustrative examples of what are more generally referred to as “processor-readable storage media” storing executable program code of one or more software programs.
Articles of manufacture comprising such processor-readable storage media are considered illustrative embodiments. A given such article of manufacture may comprise, for example, a storage array, a storage disk or an integrated circuit containing RAM, ROM, flash memory or other electronic memory, or any of a wide variety of other types of computer program products. The term “article of manufacture” as used herein should be understood to exclude transitory, propagating signals. Numerous other types of computer program products comprising processor-readable storage media can be used.
702 1 714 704 Also included in the processing device-is network interface circuitry, which is used to interface the processing device with the networkand other system components, and may comprise conventional transceivers.
702 700 702 1 The other processing devicesof the processing platformare assumed to be configured in a manner similar to that shown for processing device-in the figure.
700 100 Again, the particular processing platformshown in the figure is presented by way of example only, and systemmay include additional or alternative processing platforms, as well as numerous distinct processing platforms in any combination, with each such platform comprising one or more computers, servers, storage devices or other processing devices.
For example, other processing platforms used to implement illustrative embodiments can comprise various arrangements of converged infrastructure.
It should therefore be understood that in other embodiments different arrangements of additional or alternative elements may be used. At least a subset of these elements may be collectively implemented on a common processing platform, or each such element may be implemented on a separate processing platform.
As indicated previously, components of an information processing system as disclosed herein can be implemented at least in part in the form of one or more software programs stored in memory and executed by a processor of a processing device. For example, at least portions of the fault domain identifier processing functionality provided by one or more components of a storage system as disclosed herein are illustratively implemented in the form of software running on one or more processing devices.
It should again be emphasized that the above-described embodiments are presented for purposes of illustration only. Many variations and other alternative embodiments may be used. For example, the disclosed techniques are applicable to a wide variety of other types of information processing systems, hosts, storage systems, storage nodes, storage devices, storage processors, initiators, targets, path selection logic instances, fault domain identifier processing logic instances and other components. Also, the particular configurations of system and device elements and associated processing operations illustratively shown in the drawings can be varied in other embodiments. Moreover, the various assumptions made above in the course of describing the illustrative embodiments should also be viewed as exemplary rather than as requirements or limitations of the disclosure. Numerous other alternative embodiments within the scope of the appended claims will be readily apparent to those skilled in the art.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
February 28, 2025
September 3, 2026
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.