I'm still looking for insights into optimizing storage costs for large language model training workloads. Many of the discussions have focused on best practices, but I haven't found concrete guidance on cost-effective storage tiers, data layout, or tiered caching approaches. If anyone has experience or resources on this topic, I would greatly appreciate your advice. Thank you!