IBM i Global

 View Only

 Booting from SAN through VIOS

Magne Stamnes's profile image
Magne Stamnes posted 02/24/26 09:31 AM

I have a question for you Guru's.

I have a P9 running VIOS and an IBM i LPAR booting from SAN.

On the SAN, this LPAR has 3 hosts and each host have mapped 100 LUNs using NPIV. 300 LUNs in total.

This is working fine.

Now, I am in progress of moving this to new HW. New P11, VIOS, IBM i LPAR.

The new SAN will have 3 hosts like the old. And each host will have the same 100 LUNs mapped.

First time booting the new LPAR on the new HW, my expreience is that it will take a long time before the LPAR finds it's Load Source disk and boots.

Is there a way to shorten this time? I think the system is scanning all 100 LUNs to find the one LS.

Can I map the LS first and single, boot in manual, then from DST, shutdown the system and then map all the rest 299 LUN's?

Or will that most likely break the system in some way?

Thanks in advance!

Best regards

Magne Stamnes

Embriq AS

Norway

Satid S's profile image
Satid S

Dear Magne

Within an IBM i LPAR's allocated HW resource (virtual or otherwise) info, you use HMC to do Load Source Tagging to the particular vFC that you know connects to the Load Source Disk Unit (LUN).  This way, the search for the load source LUN starts at that vFC being tagged. 

And here just in case this is not the case for you. In the past, I used to encounter a long IPL time for many IBM i LPARs that connected to IBM Storwize SAN with NPIV which my customer perceived as abnormal because other LPARs took much less time for IPL.  What I did was checking the progress of IPL process from HMC by looking at SRC code progress of an IBM i LPAR at issue. I noticed that a particular SRC code (can't remember the code) took a long time (about 5 minutes) to change to the next SRC code. I Googled that particular SRC code and found that it indicted a step called "DASD Check" (DASD is an old IBM jargon for disk and I scratched my head why such an old jargon was still used). 

After checking around for possibilities relating to disk resource check at IBM i IPL, I found that the customer used only 2 paths for NPIV connection from IBM i LPAR to SAN.  With a look at the HW resources info in IBM i SST menu, I found many "failed" (not sure if this was the exact wording) resources that were NPIV multipath connections and suspected it might be the cause IPL process wasted time waiting for response on those failed multipath links.

Luckily I found an IBM Technote on using SST for "Reducing or removing paths for a multipath LUN" - https://www.ibm.com/support/pages/reducing-or-removing-paths-multipath-lun - (there is another way not using SST: Cloning IBM i: Resetting Multipath at  https://blog.faq400.com/en/system-administration-en/cloning-ibm-i-resetting-multipath/  ) and the IPL time became much faster. That DASD Check SRC no longer took minutes to run.  On further research, it turned out that IBM i normally established 4 paths to connect to that IBM i Storwize SAN but somehow the customer wanted to reduce to 2 paths only but was not aware to remove the unused multipath links.  

Magne Stamnes's profile image
Magne Stamnes

Hi Satid

Thanks for the answer. It did not help me but that can be because I did not give all the details.

I have tagged the vFC in the lpar profile. 

Before performing the migration from old to new HW, I created a test LUN that I attached to the host on the SAN. I booted the lpar from DVD and installed full OS with PTFs on this test LUN. Just so I could perform HW tests, like network and other things.

The migration from old to new HW was done by connecting the old and new SAN and copying all the LUNs from the old to the new SAN.

The old lpar was then shut down to allow for all data to be copied.

On the new SAN, all the copied LUNs was then attached to their respective hosts. After breaking the copy process.

The test LUN was of course removed from the host on the SAN.

It is at this point, when I am starting the new lpar, on new HW, in Manual mode from B, that the IPL takes a long time to complete. And the HMC is showing C2003160 for a long time and several times. This code is changing breefly before it returns.

It seems to me that the system is searching for the Load Source disk among the 100 LUNS attached to the tagged vFC.

And yes, the system has redundant paths to the SAN. When IBM i is up and running there is 4 active and 4 passive paths shown in SST.

And also, at this time, I cannot perform any DST options like removing old paths, because the system is starting for the first time on new HW and original disks copied from the old SAN.

I am looking for a way to shorten the time for the system to find it's Load Source disk. If it is possible.

Best regards

Magne Stamnes

Bartlomiej Grabowski's profile image
Bartlomiej Grabowski IBM Champions

Interesting question Magne, I certainly know what kind delay you mean but I don't think it is related to number of disks. . The LPAR waits in C600x code for about 10-15 min, right? I think it is related to multi-path checks and HW discovery.

Certainly there is no write IO operation to the disk if you go to DST only. I have written an article about it in my blog- https://theibmi.org/2019/08/05/recovering-an-unconfigured-load-source-disk/. 

I was   in position that not all disks were "attached" to the LPAR due to wrong zoning and data was not affected. The system would complain that some disks are missing and the only available option is go to DST. 

You can try to attach only the Load Source disk , and try to boot on P11, but I think you still will face 10-15 min delay. I think it is related to HW configuration change, and HW recovery. 

If there is an option to reduce this delay window thru some SST macro, I don't know. 

Pedro Lázaro Martínez's profile image
Pedro Lázaro Martínez

Hi.

If the problem isn't related resetting the multipath or missing disks, as Satid and Bartlomiej suggest, it might be, as you suspect, that the system remembers the original load source from the test, which is no longer connected.

In that case, I suggest you try the "Update VPD Data (Vital Product Data)" procedure in DST. VPD marks the load disk for the system for subsequent IPLs and was a common procedure back in the days of hardware migrations with internal hard drives.

In example: https://www.ibm.com/docs/en/i/7.5.0?topic=rlic-recovering-vital-product-data-information-if-partition-does-not-ipl-in-mode-b-mode

I hope this helps.

Regards

Satid S's profile image
Satid S

Dear Magne

As Mr. Bart mentioned, for the very first IBM i IPL from a new Load Source disk in an LPAR, IBM i scans all HW resources allocated to the LPAR which includes all disk units (LUNs) but not for the reason you mentioned.  There is no way to avoid this but I also notice that if the LPAR is allocated less than 0.1 processing unit CPU power or assigned too many Virtual Processor for the LPAR at less than 1 processing unit, this can contribute to the length of time it takes for the IPL.  If you allocate this small CPU power to the  LPAR, you may want to increase it a bit just for this IPL or set the LPAR as an "Uncapped" partition.   And if the LPAR is allocated 1 processing unit or less, DO NOT assign more than 1 Virtual Processor to the LPAR. 

 

Magne Stamnes's profile image
Magne Stamnes

Thanks for all the answers.

It looks to me that the answer from Pedro is most suitable for me and I will test that.

I tried this yesterday;

I have a system with 60 LUNS that I can perform tests on.

I attached the LS as single LUN to a test lpar and booted from B M.

It got past the  C2003160 quickly but got stuck on A6XX0266.

After waiting about 1 hour I attached the rest of the LUNs to the host on the SAN and waited for 1 hour again.

The same A6XX0266 was the only reaction.

So I performed a shut down of the partition from the HMC (Since I never got into DST) and tried to start the lpar again.

Now with all 60 LUNs attached.

I left the system over night and this morning the same A6XX0266 code was displayed.

So that did not work. It looks like I made it worse by first attaching only the LS.

I will do more tests today, based on Pedro's answer and I will report back.

Thanks again!

Best regards

Magne Stamnes

Magne Stamnes's profile image
Magne Stamnes

That was a quick test that failed.

I booted the lpar in D M and tried to update VPD as Pedros instructions.

The lpar claims it cannot find a valid LS.

I might have "broken" these LUN's.

I'll check to see if I can create a new copy of the LUNs from a working system so that I can test again, because if this works, then it will save me a lot of time when we actually are going to move the next system.

Regarding som other statements that you have asked:

This lpar that I am testing on, has 9 active processors and 3 TB memory, so it should be enough resources allocated.

I will also try Bartlomiej's suggestion and report back.

Magne Stamnes's profile image
Magne Stamnes

I just tried Bartlomiej's idea by trying to change the data on the disk but when I hit F9 to save I was told that the save did not work.

The change was not saved.

So my "problem" with these LUNs is different.

I need to make a new copy and test more.

Magne

Pedro Lázaro Martínez's profile image
Pedro Lázaro Martínez

Hi Magne.

I'm sorry to hear it didn't work.

Using the HMC, you specify which vFC IOA the boot disk should be searched for in (Tagged IO settings).

But the VPD marks exactly which disk is the LS, and the system remembers that information for subsequent IPLs.

This doesn't prevent the system from ALWAYS checking the integrity of the ASP (DADS check), and regarding that, the A6xx0255 codes you're seeing are specific to missing disks in the ASP. Perhaps there are volumes missing from the storage or the copy is corrupted.

For IBM i, it only matters that ALL the disks in the ASP are present, regardless of which resource they connect to or how they are grouped. If all the disks in the ASP are present and the boot process is tagged in the HMC, the system should perform an IPL. The only remaining issue is that the LS serial number may have changed compared to a previous IPL, and this can be corrected, if necessary, during the VPD update procedure.

If VPD procedure doesn't work, may be because  it can't actually find a loading disk to propose for marking.

However, based on your last comment, it seems the disk copy is incorrect. In similar migrations, when the storage system changes in addition to the server, I usually use the virtualization or image-mode volume migration capabilities of the storage system (I assume the storage is IBM Flash System). It's safe and fast, resulting in smooth migrations.

Check the volume copy. Either disks are missing from the ASP or the copy is incorrect and the volume contents are corrupted and unusable.

Please let us know the result.

Regards.

Bartlomiej Grabowski's profile image
Bartlomiej Grabowski IBM Champions

Magne,

According to your latest updated, you are facing completely different problem which we thought. 

I have migrated dozen of LPARs from P9 to P11 with hundreds of LUNs and I have never waited more than 30 min. If you wait for the entire night, there is definitely something wrong. 

Maybe few questions first,

  • Do you use an external storage connected thru NPiV ?
  • Can you share type and model external storage used?
  • I assume during migration you take the same set of LUNs from the same storage and rezone to new P11 LPAR. Change SAN zoning (on Fabrics), change hostconnect on the storage device (P11 LPAR comes with new WWPN which need to be updated on Fabrics and storage device). 
  • I assume if you perform an IPL on P9 , the process takes maximum few minutes, the LPAR does not go to A6XX0266 during boot process. 
  • A6XX0266 shows when an LPAR lost connection with a storage device.  It can happen during IPL, but only if someone at the same time play with storage configuration. 
  • Have you migrated any LPAR to P11 already, Is it booting normally? 

The only thing which comes to my mind is some incompatibility between SAN and the SLIC. The situation which you are facing is not normal. 

Magne Stamnes's profile image
Magne Stamnes

Hi Bartlomiej.

We have already migrated on lpar from the old HW to the new HW. That system was a test lpar with 61 LUN's.

We copied the LUNs from the old SAN to the new SAN using SAN technology.

The startup of the new test lpar used aprox 30 minutes to start. If the number of LUNs attached does not matter, then I beleive I can wait it out on the next lpar we are going to migrate. The next lpar has about 270 LUNs.

The old SAN is IBM Flash 900 behind IBM SVC. The new SAN is IBM Flash 9500.

We use NPIV.

Bartlomiej Grabowski's profile image
Bartlomiej Grabowski IBM Champions

To be more precise - confirm or deny. 

You create same number of LUNs with the same size on IBM Flash950  as exists on SVC ( I assume disks to LPAR are mapped on SVC layer)

  • Configuration replication between SVC and Flash950
  • Start the replication (asynchronous) ,
  • Power of the LPAR on P9
  • Wait for the replication to process all the backlog
  • Stop and break the replication, 
  • Power On the LPAR on P11

Is this the way how are you doing?

 

Pedro Lázaro Martínez's profile image
Pedro Lázaro Martínez

Hi.

Perhaps Bartlomiej hit the nail on the head when he mentioned the LIC.
When migrating from POWER9 to POWER10 or POWER11, the IBM i operating system must not only be a version compatible with the new hardware but also have the minimum LIC "resave" required for the new POWER server.

If it doesn't, you'll have problems with IPL at microcode level. You must ensure the current LIC resave level by checking the RExxxxx marker PTF level of microcode 5770999:

https://www.ibm.com/support/pages/ibm-i-resaves

If it's not adequate, before IPL on the new server, you'll need to perform a slip install of the microcode, in current POWER9, to an updated resave for the same release, as well as load the latest PTFs, including the hardware PTF group (just in case), and the latest Technology Refresh if a requisite for the installed LIC resave.

This wouldn't be the first time a migration has failed for this reason, based on the assumption that simply having the IBMi release supported by the new server is sufficient.

Regards.

Magne Stamnes's profile image
Magne Stamnes

Bartlomiej

I can confirm that this is what we are doing. In that order.

And this worked fine for our test server with 61 LUNs. And since I am afraid that our 30 min wait time will increase 3 fold when the number of LUNS on the next move is much higher.

Pedro: The first system, the test system, was migrated without a problem. So the LIC should not be a problem or out of date.

The next system that we plan to move has a lower CUM and group PTF levels than the test system but they should be fresh enough.

Now; Please note that I am trying to find a way to not being forced to wait out the system search for HW. I'm trying to find a way that shortens the time for the system to boot.

But if that is not possible, I will wait while the system is doing what the system needs to do before it boots.

Bartlomiej Grabowski's profile image
Bartlomiej Grabowski IBM Champions

Check Technology Level - this PTF group defines what HW is supported. 

The only thing what I can advise (but don't know if feasible in your environment) . Perform an IPL on P9 from Flash950 storage. 

You are doing two operations at the same time (storage migration and HW migration) in one step. It is hard to say which element goes wrong. 

Magne Stamnes's profile image
Magne Stamnes

Hello guys!

I performed the following tests but I feel that they are not conclusive.

First I made a Snapshot of the already migrated system with 61 LUNs. I mapped the Snapshot LUNS to a new partition (The partition exists and have been tested, but is currently offline) and I booted from D M.

This was quite quick and in DST I performed the update VPD. The system had found the LS.

I then performed a shutdown from DST (Operator panel function F10) but that failed so I had do shut it down from HMC. (Don't remember the SRC code)

Then I booted from B M and the system came up rather quickly. All disks present.

I then Powered off the system, deleted the Snapshot, created a new Snapshot and mapped the disks again.

Booted from B M this time and this time it took longer to go into DST.  I will guess around 10 minutes. 

It used those minutes on C2003160.

But the real test would have been to copy LUNS from the old to the new SAN once more, and test on those. That would be more conclusive I guess.

But since we are going to move the next lpar on Saturday, I will just be patient and wait until the system finds the disks by itselves. I don't want to break something based on this test.

However, it is interesting.  I might test some more later.

Thanks anyway for all your good tips. I appreciate it.

Pedro Lázaro Martínez's profile image
Pedro Lázaro Martínez

Hi.

I hope you can resolve the issue satisfactorily.

Regardless of the disk contents (which are obviously identical in the snapshots), the system treats them as new disks since they display different UIDs and paths compared to the originals.

Therefore, it will always spend time checking the DASD configuration and the integrity of the ASP, ensuring that all disks report.


Furthermore, after the first IPL, a MULTIPATHRESSETER should be performed to prevent subsequent IPLs from wasting time rechecking paths that no longer exist. This minimizes the time required for subsequent boot times.


This is normal and expected behavior when performing the first IPL from cloned disks, whether from a snapshot or replication. It is typical in Full System Replication or flashcopy partition scenarios. If the POWER server is physically different, then additional time is spent because Activation Engine tasks (SRC codes such as C600450A and others with the letters *AE* are displayed).

But... this is IBMi. IBMi always ends up booting!

Regards

Satid S's profile image
Satid S

Dear Magne

I Google "ibm i src a6000266" and find that your problem is a known BUG (since around April last year) as described in this IBM Technote: A6000266 when adding virtual disk to ASP on i client partition hosted by VIOS. at  https://www.ibm.com/support/pages/a6000266-when-adding-virtual-disk-asp-i-client-partition-hosted-vios      You should apply PTFs that fix this issue and you will find the PTF IDs from the Google search result named Known Issue DT30NNNN (there are multiple entries of this, so you choose the one that applies to your IBM i release). 

 

Magne Stamnes's profile image
Magne Stamnes

Thank you Pedro for all your valuable input.

In my migration plan, I have a step for MULTIPATHRESSETER and also to remove missing HW from SST/DST.

Satid; The a6000266 came because I tried to add the rest of the LUNs to the lpar after starting it from the known LS disk as a single drive. This will not happen during the migration.

But thank you for the input :-)

Have a great weekend all of you gurus out there!

/Magne