[<prev] [next>] [thread-next>] [day] [month] [year] [list]
Message-Id: <20240312061729.1997111-1-horenchuang@bytedance.com>
Date: Tue, 12 Mar 2024 06:17:26 +0000
From: "Ho-Ren (Jack) Chuang" <horenchuang@...edance.com>
To: "Gregory Price" <gourry.memverge@...il.com>,
aneesh.kumar@...ux.ibm.com,
mhocko@...e.com,
tj@...nel.org,
john@...alactic.com,
"Eishan Mirakhur" <emirakhur@...ron.com>,
"Vinicius Tavares Petrucci" <vtavarespetr@...ron.com>,
"Ravis OpenSrc" <Ravis.OpenSrc@...ron.com>,
"Alistair Popple" <apopple@...dia.com>,
"Rafael J. Wysocki" <rafael@...nel.org>,
Len Brown <lenb@...nel.org>,
Dan Williams <dan.j.williams@...el.com>,
Vishal Verma <vishal.l.verma@...el.com>,
Dave Jiang <dave.jiang@...el.com>,
Andrew Morton <akpm@...ux-foundation.org>,
Jonathan Cameron <Jonathan.Cameron@...wei.com>,
Huang Ying <ying.huang@...el.com>,
"Ho-Ren (Jack) Chuang" <horenchuang@...edance.com>,
linux-acpi@...r.kernel.org,
linux-kernel@...r.kernel.org,
nvdimm@...ts.linux.dev,
linux-cxl@...r.kernel.org,
linux-mm@...ck.org
Cc: "Ho-Ren (Jack) Chuang" <horenc@...edu>,
"Ho-Ren (Jack) Chuang" <horenchuang@...il.com>,
qemu-devel@...gnu.org
Subject: [PATCH v2 0/1] Improved Memory Tier Creation for CPUless NUMA Nodes
When a memory device, such as CXL1.1 type3 memory, is emulated as
normal memory (E820_TYPE_RAM), the memory device is indistinguishable
from normal DRAM in terms of memory tiering with the current implementation.
The current memory tiering assigns all detected normal memory nodes
to the same DRAM tier. This results in normal memory devices with
different attributions being unable to be assigned to the correct memory tier,
leading to the inability to migrate pages between different types of memory.
https://lore.kernel.org/linux-mm/PH0PR08MB7955E9F08CCB64F23963B5C3A860A@PH0PR08MB7955.namprd08.prod.outlook.com/T/
This patchset automatically resolves the issues. It delays the initialization
of memory tiers for CPUless NUMA nodes until they obtain HMAT information
at boot time, eliminating the need for user intervention.
If no HMAT is specified, it falls back to using `default_dram_type`.
Example usecase:
We have CXL memory on the host, and we create VMs with a new system memory
device backed by host CXL memory. We inject CXL memory performance attributes
through QEMU, and the guest now sees memory nodes with performance attributes
in HMAT. With this change, we enable the guest kernel to construct
the correct memory tiering for the memory nodes.
-v2:
Thanks to Ying's comments,
* Rewrite cover letter & patch description
* Rename functions, don't use _hmat
* Abstract common functions into find_alloc_memory_type()
* Use the expected way to use set_node_memory_tier instead of modifying it
-v1:
* https://lore.kernel.org/linux-mm/20240301082248.3456086-1-horenchuang@bytedance.com/T/
Ho-Ren (Jack) Chuang (1):
memory tier: acpi/hmat: create CPUless memory tiers after obtaining
HMAT info
drivers/acpi/numa/hmat.c | 11 ++++++
drivers/dax/kmem.c | 13 +------
include/linux/acpi.h | 6 ++++
include/linux/memory-tiers.h | 8 +++++
mm/memory-tiers.c | 70 +++++++++++++++++++++++++++++++++---
5 files changed, 92 insertions(+), 16 deletions(-)
--
Ho-Ren (Jack) Chuang
Powered by blists - more mailing lists