Hi,
We have an SVC cluster (8.7.0.7 code) made up with 4 nodes, 2 x SV2 and 2 x SV3. Backend storage are 2 x FS7300's (8.7.0.7 code). Came in this morning on Monday and we had a 1220 error on the SVC, "Remote port excluded" from Friday daytime. When going through the Fix procedure on the SVC, it comes back on the 2nd screen, after clicking Next on the first screen, with message saying it can'f fix the problem:
"The system is unable to proceed with the fix procedure, click Close to return to the management GUI"
I also had a number of "Login excluded - 1230", "Login transport fault - 1360" errors, and "Number of device logins reduced - 1630". There were also a number of SCSI ERP errors which were monitoring and have expired and fixed themselves after 24 hours. All the additional errors occurred Friday night and early hours of Saturday morning, and none since. I have marked all errors as Fixed, as this was available, in the hope this would re-set the degraded state of a single mdisk, along with running Discover storage repeatedly.
This has left the SVC reported degraded connections on "Logical Components" as reported on the SVC Dashboard:
- 2 degraded "External Storage Controllers", these are the two controllers from one of the FS7300's
- 1 degraded "External Mdisks", just the one mdisk degraded with 12 paths instead of the usual 16 that all other mdisks have
- 1 degraded "Pools", i.e. the pool which has the mdisk issue
So we have a pool with 51 mdisks in it, from a single FS7300, and only one of the mdisks is reporting degraded with 12 paths instead of the usual 16.
Looking on the SAN at porterrshow I could see there were errors on one of the FC connections on one of the SVC nodes. The porterrshow gets cleared automatically by scripts in the early hours of Monday this morning, and since it was cleared, we have no further port errors. On the backend FS7300's we have no errors reported at all, nothing degraded, all green and online in the GUI's.
So where do I go from here to try and get the single mdisk from a single backend array out of Degraded and back to Online status with a full 16 paths? My suggestion would be to disable the original SVC port which had the errors, wait 15 mins for the SVC to register paths offline, forcing all the mdisks back to 12 paths, and then re-enable the SVC port to hopefully force all mdisks back to 16 paths. Not 100% sure if this would work though.
I can raise an incident with IBM, but with my customer, I cannot send logs to IBM, annoying I know, but that's the position I am in. But I thought I would try here in case anyone might have seen this 1220 Remote port excluded error before and how to get the single mdisk back to full path redundancy.
Thanks, Andy