Skip to content

test: describe hardware scenarios in YAML, play them on the adapters - #102

Open
g-carre wants to merge 4 commits into
mainfrom
test/ARTESCA-17960-megaraid-disk-failure-scenario
Open

g-carre wants to merge 4 commits into
mainfrom
test/ARTESCA-17960-megaraid-disk-failure-scenario

Conversation

@g-carre

@g-carre g-carre commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Refs ARTESCA-17960 / RD-2154. Used by scality/disk-management-agent#74.

Adds hardware scenarios written in YAML. Each scenario takes a real controller capture, then changes drives and volumes phase by phase, and gives the table of disks expected in each phase. The scenarios are played through the raidmgmt adapters here, and through the agent in #74.

The scenario: scenario/scenarios/megaraid-disk-failure.yaml

- name: failed
  description: |
    251:3 fails. Its RAID0 volume goes offline and is no longer a block
    device. The other drives must keep their paths.
  knownBug: ARTESCA-17960
  changes:
    drives:
      "251:3": {state: Failed, serial: SIMD0003}
    volumes:
      "237": {state: OfLn, exposed: false}
  expect: |
    slot    status  serial                devicePath  permanentPath
    251:1   Used    SIMD0001000000000000  /dev/sdl    /dev/disk/by-id/wwn-0x600062b2…0239
    251:3   Failed  SIMD0003              -           -
    ...

The phases are healthy, then failed (251:3 fails and its RAID0 goes offline), then replaced (a new drive, UGood), then restored (a new RAID0 volume). The capture is real storcli output of the RD-2154 lab node, kept to its first 4 drives (251:1-4) to keep the tables short, with serials, WWNs and controller identifiers replaced by fake ones. The failure is derived from what that node showed when 251:3 failed, as explained in the scenario description.

Known bugs are expected failures

A phase with knownBug must differ from its table. The test passes and logs the differences:

expected failure, ARTESCA-17960:
    251:1 devicePath: got -, want /dev/sdl
    ... (the healthy drives lose their paths)
    LogicalVolumes: reported but not expected: failed to get logical volumes: failed to get logical
    volume 237: failed to get paths: failed to compute paths from physical drive: failed to compute permanent path

Once the phase matches its table, the test fails with the bug looks fixed, remove its knownBug line, so the fix PR has to update the scenario. This is also the result on the #91 branch: #91 only handles volumes with several drives, and VD 237 is a single-drive RAID0.

Layout: a separate module, github.com/scality/raidmgmt/scenario

  • scenarios/*.yaml: the scenarios.
  • scenario: loads them and compares the reported disks with a table, including the known-bug rule.
  • megaraidsim: answers storcli64 commands from a capture with the changes of a phase applied, and shows the /dev/disk/by-id links of the exposed volumes.
  • megaraidsim/adapter_test.go: plays every MegaRAID scenario through this repository's megaraid adapter, thanks to a replace ../.

It is a separate module so that consumers can play the scenarios without changing the raidmgmt version they test. The agent stays on v0.16.0, and even the 4.3 agent (raidmgmt v0.14.0) could use it. Adding a controller means a capture plus a backend like megaraidsim; adding a scenario on an existing capture is just a YAML file. See scenario/README.md.

CI

  • The module list now includes nested modules, and tests run in each module's directory.
  • The scenarios run in the normal tests job. There is no build tag and no extra job.
  • make tests also runs them.

@g-carre
g-carre requested a review from a team as a code owner October 9, 2026 15:53
Comment thread pkg/implementation/raidcontroller/megaraid/scenario_diskfailure_test.go Outdated
Comment thread .github/workflows/pre-merge.yaml Outdated
Comment thread scenario/scenario.go Outdated
@g-carre g-carre changed the title test(megaraid): replay a disk failure scenario from lab captures test: describe hardware scenarios in YAML, play them on the adapters Oct 9, 2026
Comment thread scenario/scenario.go
Comment thread Makefile
Comment thread scenario/scenario.go Outdated
@g-carre
g-carre force-pushed the test/ARTESCA-17960-megaraid-disk-failure-scenario branch from 920e9b3 to 5b6c435 Compare October 9, 2026 17:30
Comment thread scenario/megaraidsim/megaraidsim.go
@g-carre
g-carre force-pushed the test/ARTESCA-17960-megaraid-disk-failure-scenario branch from ccebc24 to c758ee1 Compare October 9, 2026 17:35
Comment thread scenario/megaraidsim/captures/paris4-node1/c0.json Outdated
Comment thread scenario/scenario.go
@g-carre
g-carre force-pushed the test/ARTESCA-17960-megaraid-disk-failure-scenario branch 2 times, most recently from 5bbed0e to 10340b5 Compare October 9, 2026 17:37
Comment thread scenario/megaraidsim/megaraidsim.go
Comment thread scenario/scenario.go Outdated
@g-carre
g-carre force-pushed the test/ARTESCA-17960-megaraid-disk-failure-scenario branch from 10340b5 to 2b94d50 Compare October 9, 2026 17:44
Add github.com/scality/raidmgmt/scenario, a separate module so that
consumers can play the scenarios without changing the raidmgmt version they
test:

- scenarios/*.yaml: a capture of a real controller, then phases changing
  drives and volumes, each with the table of disks expected and, when a bug
  makes it differ today, the ticket and the table reported because of it
- scenario: loads scenarios strictly, rejecting unknown keys and phases
  without table, and compares what is reported with a table. A phase with a
  known bug passes only while the table of the bug is reported; it fails
  once the expected table is (the bug is fixed) and when neither is (the
  behaviour changed). A drive reported twice is a difference.
- megaraidsim: answers storcli64 commands from a capture with the changes of
  a phase applied, and shows the /dev/disk/by-id links of exposed volumes
- megaraidsim/adapter_test.go: plays every MegaRAID scenario through the
  megaraid adapter of this repository

The first scenario, megaraid-disk-failure, replays real storcli output of a
lab node, kept to its first 4 single-drive RAID0 volumes and with fake
serials and WWNs, then drive 251:3 failing, being replaced and getting a new
volume. Its failed phase documents ARTESCA-17960: the offline volume of the
failed drive makes LogicalVolumes fail for the whole controller, so no drive
gets a path.

CI and the release gate now test and lint every module.

Refs: ARTESCA-17960
@g-carre
g-carre force-pushed the test/ARTESCA-17960-megaraid-disk-failure-scenario branch from 2b94d50 to 5201052 Compare October 9, 2026 17:44
@g-carre
g-carre force-pushed the test/ARTESCA-17960-megaraid-disk-failure-scenario branch from 14ede3a to 1f7e333 Compare October 9, 2026 19:15
Comment thread scenario/megaraidsim/megaraidsim.go Outdated
Comment thread scenario/storcli2sim/storcli2sim.go Outdated
Add storcli2sim, the backend of storcli2/perccli2 controllers: it answers the
commands of the raidmgmt storcli2 adapter (controller listing, drives and
volumes) from a capture with the changes of a phase applied. The capture is
the storcli2 test data of this repository, kept to 4 RAID0 drives, with fake
serials, WWNs, SAS addresses, volume and controller identifiers.

storcli2-disk-failure plays the megaraid-disk-failure story on it. The
healthy drives keep their paths when 306:2 fails, but the offline volume of
the failed drive is still reported with a permanent path built from its NAA
Id, although it has no block device: documented as a known bug of
ARTESCA-17960.

Drive changes take a status, as storcli2 reports it apart from the state.
scenario.Report collects the disks an adapter reports, shared by the
megaraid and storcli2 adapter tests; like disk-management-agent, it lists the
controllers first.

Refs: ARTESCA-17960
@g-carre
g-carre force-pushed the test/ARTESCA-17960-megaraid-disk-failure-scenario branch from 1f7e333 to 2e737a1 Compare October 9, 2026 19:21
Comment thread scenario/ssaclisim/ssaclisim.go Outdated
…ollers

Add ssaclisim, the backend of ssacli controllers: it answers the commands of
the raidmgmt ssacli getters (controller listing, physical and logical drive
details, show config) from a capture with the changes of a phase applied.
ssacli prints text, so the backend edits the lines of the captured sections:
drive and logical drive status, serials, disk names, and moves a drive to
the Unassigned section once its array is gone. The capture is the ssacli
test data of this repository, kept to 4 single-drive RAID0 arrays, with fake
serials, WWIDs and identifiers.

ssacli-disk-failure plays the megaraid-disk-failure story on it. The ssacli
getters already report it as expected: the healthy drives keep their paths,
the failed one is Failed, and the replacement drive is UnassignedGood.

Refs: ARTESCA-17960
@g-carre
g-carre force-pushed the test/ARTESCA-17960-megaraid-disk-failure-scenario branch from e7d17c7 to ce633e6 Compare October 9, 2026 21:41
Comment thread scenario/megaraidsim/megaraidsim.go Outdated
Comment thread scenario/storcli2sim/storcli2sim.go Outdated
storcli and storcli2 print the drive group (DG) of a drive as a number, and
"-" only for an unconfigured drive. The simulators wrote a group change
as a string, so a phase replayed output real controllers never print.
scenario.NumericGroup converts a group to a number, keeps "-", and rejects
anything else.

Refs: ARTESCA-17960
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant