[<prev] [next>] [thread-next>] [day] [month] [year] [list]
Message-ID: <cover.e1ebdd6cab9bde0d232c1810deacf0bae25e6707.1732239628.git-series.apopple@nvidia.com>
Date: Fri, 22 Nov 2024 12:40:21 +1100
From: Alistair Popple <apopple@...dia.com>
To: dan.j.williams@...el.com,
linux-mm@...ck.org
Cc: Alistair Popple <apopple@...dia.com>,
lina@...hilina.net,
zhang.lyra@...il.com,
gerald.schaefer@...ux.ibm.com,
vishal.l.verma@...el.com,
dave.jiang@...el.com,
logang@...tatee.com,
bhelgaas@...gle.com,
jack@...e.cz,
jgg@...pe.ca,
catalin.marinas@....com,
will@...nel.org,
mpe@...erman.id.au,
npiggin@...il.com,
dave.hansen@...ux.intel.com,
ira.weiny@...el.com,
willy@...radead.org,
djwong@...nel.org,
tytso@....edu,
linmiaohe@...wei.com,
david@...hat.com,
peterx@...hat.com,
linux-doc@...r.kernel.org,
linux-kernel@...r.kernel.org,
linux-arm-kernel@...ts.infradead.org,
linuxppc-dev@...ts.ozlabs.org,
nvdimm@...ts.linux.dev,
linux-cxl@...r.kernel.org,
linux-fsdevel@...r.kernel.org,
linux-ext4@...r.kernel.org,
linux-xfs@...r.kernel.org,
jhubbard@...dia.com,
hch@....de,
david@...morbit.com
Subject: [PATCH v3 00/25] fs/dax: Fix ZONE_DEVICE page reference counts
Main updates since v2:
- Rename the DAX specific dax_insert_XXX functions to vmf_insert_XXX
and have them pass the vmf struct.
- Seperate out the device DAX changes.
- Restore the page share mapping counting and associated warnings.
- Rework truncate to require file-systems to have previously called
dax_break_layout() to remove the address space mapping for a
page. This found several bugs which are fixed by the first half of
the series. The motivation for this was initially to allow the FS
DAX page-cache mappings to hold a reference on the page.
However that turned out to be a dead-end (see the comments on patch
21), but it found several bugs and I think overall it is an
improvement so I have left it here.
Device and FS DAX pages have always maintained their own page
reference counts without following the normal rules for page reference
counting. In particular pages are considered free when the refcount
hits one rather than zero and refcounts are not added when mapping the
page.
Tracking this requires special PTE bits (PTE_DEVMAP) and a secondary
mechanism for allowing GUP to hold references on the page (see
get_dev_pagemap). However there doesn't seem to be any reason why FS
DAX pages need their own reference counting scheme.
By treating the refcounts on these pages the same way as normal pages
we can remove a lot of special checks. In particular pXd_trans_huge()
becomes the same as pXd_leaf(), although I haven't made that change
here. It also frees up a valuable SW define PTE bit on architectures
that have devmap PTE bits defined.
It also almost certainly allows further clean-up of the devmap managed
functions, but I have left that as a future improvment. It also
enables support for compound ZONE_DEVICE pages which is one of my
primary motivators for doing this work.
Signed-off-by: Alistair Popple <apopple@...dia.com>
---
Cc: lina@...hilina.net
Cc: zhang.lyra@...il.com
Cc: gerald.schaefer@...ux.ibm.com
Cc: dan.j.williams@...el.com
Cc: vishal.l.verma@...el.com
Cc: dave.jiang@...el.com
Cc: logang@...tatee.com
Cc: bhelgaas@...gle.com
Cc: jack@...e.cz
Cc: jgg@...pe.ca
Cc: catalin.marinas@....com
Cc: will@...nel.org
Cc: mpe@...erman.id.au
Cc: npiggin@...il.com
Cc: dave.hansen@...ux.intel.com
Cc: ira.weiny@...el.com
Cc: willy@...radead.org
Cc: djwong@...nel.org
Cc: tytso@....edu
Cc: linmiaohe@...wei.com
Cc: david@...hat.com
Cc: peterx@...hat.com
Cc: linux-doc@...r.kernel.org
Cc: linux-kernel@...r.kernel.org
Cc: linux-arm-kernel@...ts.infradead.org
Cc: linuxppc-dev@...ts.ozlabs.org
Cc: nvdimm@...ts.linux.dev
Cc: linux-cxl@...r.kernel.org
Cc: linux-fsdevel@...r.kernel.org
Cc: linux-mm@...ck.org
Cc: linux-ext4@...r.kernel.org
Cc: linux-xfs@...r.kernel.org
Cc: jhubbard@...dia.com
Cc: hch@....de
Cc: david@...morbit.com
Alistair Popple (25):
fuse: Fix dax truncate/punch_hole fault path
fs/dax: Return unmapped busy pages from dax_layout_busy_page_range()
fs/dax: Don't skip locked entries when scanning entries
fs/dax: Refactor wait for dax idle page
fs/dax: Create a common implementation to break DAX layouts
fs/dax: Always remove DAX page-cache entries when breaking layouts
fs/dax: Ensure all pages are idle prior to filesystem unmount
fs/dax: Remove PAGE_MAPPING_DAX_SHARED mapping flag
mm/gup.c: Remove redundant check for PCI P2PDMA page
pci/p2pdma: Don't initialise page refcount to one
mm: Allow compound zone device pages
mm/memory: Enhance insert_page_into_pte_locked() to create writable mappings
mm/memory: Add vmf_insert_page_mkwrite()
huge_memory: Allow mappings of PUD sized pages
huge_memory: Allow mappings of PMD sized pages
memremap: Add is_device_dax_page() and is_fsdax_page() helpers
gup: Don't allow FOLL_LONGTERM pinning of FS DAX pages
proc/task_mmu: Ignore ZONE_DEVICE pages
memcontrol-v1: Ignore ZONE_DEVICE pages
mm/mlock: Skip ZONE_DEVICE PMDs during mlock
fs/dax: Properly refcount fs dax pages
device/dax: Properly refcount device dax pages when mapping
mm: Remove pXX_devmap callers
mm: Remove devmap related functions and page table bits
Revert "riscv: mm: Add support for ZONE_DEVICE"
Documentation/mm/arch_pgtable_helpers.rst | 6 +-
arch/arm64/Kconfig | 1 +-
arch/arm64/include/asm/pgtable-prot.h | 1 +-
arch/arm64/include/asm/pgtable.h | 24 +-
arch/powerpc/Kconfig | 1 +-
arch/powerpc/include/asm/book3s/64/hash-4k.h | 6 +-
arch/powerpc/include/asm/book3s/64/hash-64k.h | 7 +-
arch/powerpc/include/asm/book3s/64/pgtable.h | 52 +---
arch/powerpc/include/asm/book3s/64/radix.h | 14 +-
arch/powerpc/mm/book3s64/hash_pgtable.c | 3 +-
arch/powerpc/mm/book3s64/pgtable.c | 8 +-
arch/powerpc/mm/book3s64/radix_pgtable.c | 5 +-
arch/powerpc/mm/pgtable.c | 2 +-
arch/riscv/Kconfig | 1 +-
arch/riscv/include/asm/pgtable-64.h | 20 +-
arch/riscv/include/asm/pgtable-bits.h | 1 +-
arch/riscv/include/asm/pgtable.h | 17 +-
arch/x86/Kconfig | 1 +-
arch/x86/include/asm/pgtable.h | 51 +---
arch/x86/include/asm/pgtable_types.h | 5 +-
drivers/dax/device.c | 15 +-
drivers/gpu/drm/nouveau/nouveau_dmem.c | 3 +-
drivers/nvdimm/pmem.c | 4 +-
drivers/pci/p2pdma.c | 19 +-
fs/dax.c | 354 ++++++++++++++-----
fs/ext4/inode.c | 43 +--
fs/fuse/dax.c | 35 +--
fs/fuse/virtio_fs.c | 3 +-
fs/proc/task_mmu.c | 18 +-
fs/userfaultfd.c | 2 +-
fs/xfs/xfs_inode.c | 40 +-
fs/xfs/xfs_inode.h | 3 +-
fs/xfs/xfs_super.c | 18 +-
include/linux/dax.h | 23 +-
include/linux/huge_mm.h | 22 +-
include/linux/memremap.h | 28 +-
include/linux/migrate.h | 4 +-
include/linux/mm.h | 40 +--
include/linux/mm_types.h | 14 +-
include/linux/mmzone.h | 8 +-
include/linux/page-flags.h | 6 +-
include/linux/pfn_t.h | 20 +-
include/linux/pgtable.h | 21 +-
include/linux/rmap.h | 15 +-
lib/test_hmm.c | 3 +-
mm/Kconfig | 4 +-
mm/debug_vm_pgtable.c | 59 +---
mm/gup.c | 176 +---------
mm/hmm.c | 12 +-
mm/huge_memory.c | 233 ++++++++-----
mm/internal.h | 2 +-
mm/khugepaged.c | 2 +-
mm/mapping_dirty_helpers.c | 4 +-
mm/memcontrol-v1.c | 2 +-
mm/memory-failure.c | 6 +-
mm/memory.c | 126 ++++---
mm/memremap.c | 59 +--
mm/migrate_device.c | 9 +-
mm/mlock.c | 2 +-
mm/mm_init.c | 23 +-
mm/mprotect.c | 2 +-
mm/mremap.c | 5 +-
mm/page_vma_mapped.c | 5 +-
mm/pagewalk.c | 8 +-
mm/pgtable-generic.c | 7 +-
mm/rmap.c | 49 +++-
mm/swap.c | 2 +-
mm/truncate.c | 12 +-
mm/userfaultfd.c | 5 +-
mm/vmscan.c | 5 +-
70 files changed, 886 insertions(+), 920 deletions(-)
base-commit: 81983758430957d9a5cb3333fe324fd70cf63e7e
--
git-series 0.9.1
Powered by blists - more mailing lists