mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Huang, Ying" <ying.huang@linux.alibaba.com>
To: Qiliang Yuan <odys.yuan@gmail.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	 David Hildenbrand <david@kernel.org>,  Zi Yan <ziy@nvidia.com>,
	 Matthew Brost <matthew.brost@intel.com>,
	 Joshua Hahn <joshua.hahnjy@gmail.com>,
	Byungchul Park <byungchul@sk.com>,
	 Gregory Price <gourry@gourry.net>,
	Alistair Popple <apopple@nvidia.com>,
	 linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array()
Date: Mon, 05 Oct 2026 21:24:07 +0800	[thread overview]
Message-ID: <87v77ggu14.fsf@DESKTOP-5N7EMDA> (raw)
In-Reply-To: <20261002-bug-mm-move-pages-stat-batch-v2-1-f73b5d20519f@gmail.com> (Qiliang Yuan's message of "Fri, 02 Oct 2026 09:25:26 +0800")

Hi, Qiliang,

Qiliang Yuan <odys.yuan@gmail.com> writes:

> move_pages() with a NULL node list reports the node of each page. RDMA
> and KV-cache transfer engines use it to find where large registered
> buffers live, querying every 4K page of buffers that span hundreds of
> gigabytes.
>
> do_pages_stat_array() looks up the VMA and walks the page tables from
> the top for every address, taking and dropping the PTE lock each time.
> That costs about 105 ns per page, so a 16 GiB buffer takes 486 ms.
>
> Callers almost always pass consecutive addresses. Group them into runs
> and walk each run with walk_page_range(), which looks up each VMA and
> PTE table once and answers every page under it while holding the lock.
> Report pages as folio_walk_start() with FW_ZEROPAGE found them: the node
> of a normal folio, -EFAULT for the zero page or an address outside any
> VMA, and -ENOENT otherwise. Handle PUD and hugetlb leaves in their own
> callbacks so that the walk never splits them.
>
> On 7.3-rc5 in a 16-vCPU VM, querying every page of a populated 4 GiB
> buffer:
>
>                 before    after
>   4K pages      105 ns    31.9 ns
>   THP           90 ns     23.1 ns

As pointed out by David, nanosecond-level optimization for a not-so-hot
path isn't very attractive.  I understand the target of your
optimization is not the performance of a single page but that of a large
number of pages (such as 16 GiB).  So, please describe more clearly why
your change is necessary, for example, by providing the performance
improvement of querying 16 GiB memory.

Additionally, the raw performance number depends on the system under
test.  Please provide a little more information about your testing
system, for example, the CPU architecture, generation, physical core
count, etc.  For comparison, the performance improvement percentage
would also be helpful.

---
Best Regards,
Huang, Ying

> Suggested-by: Zi Yan <ziy@nvidia.com>
> Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
> ---
>  mm/migrate.c | 188 ++++++++++++++++++++++++++++++++++++++++++++++++++---------
>  1 file changed, 162 insertions(+), 26 deletions(-)
>
> diff --git a/mm/migrate.c b/mm/migrate.c
> index 15b45832bcfa7..f4d8bfb9b7b4a 100644
> --- a/mm/migrate.c
> +++ b/mm/migrate.c
> @@ -2451,44 +2451,180 @@ static int do_pages_move(struct mm_struct *mm, nodemask_t task_nodes,
>  	return err;
>  }
>  
> +struct pages_stat_walk {
> +	unsigned long start;
> +	int *status;
> +};
> +
> +static void pages_stat_set(struct pages_stat_walk *psw, unsigned long addr,
> +			   unsigned long end, int stat)
> +{
> +	int *status = psw->status + ((addr - psw->start) >> PAGE_SHIFT);
> +
> +	/* end wraps to 0 for the last page of the address space */
> +	for (; addr != end; addr += PAGE_SIZE)
> +		*status++ = stat;
> +}
> +
> +static int folio_stat(struct folio *folio)
> +{
> +	if (is_zero_folio(folio) || is_huge_zero_folio(folio))
> +		return -EFAULT;
> +	if (folio_is_zone_device(folio))
> +		return -ENOENT;
> +	return folio_nid(folio);
> +}
> +
> +/* Report pages the same way folio_walk_start() with FW_ZEROPAGE finds them. */
> +static int pages_stat_pud_entry(pud_t *pudp, unsigned long addr,
> +				unsigned long end, struct mm_walk *walk)
> +{
> +	struct page *page;
> +	spinlock_t *ptl;
> +	pud_t pud;
> +	int stat;
> +
> +	if (!IS_ENABLED(CONFIG_PGTABLE_HAS_HUGE_LEAVES))
> +		return 0;
> +	pud = pudp_get(pudp);
> +	if (pud_present(pud) && !pud_leaf(pud))
> +		return 0;
> +
> +	ptl = pud_lock(walk->mm, pudp);
> +	pud = pudp_get(pudp);
> +	if (pud_present(pud) && !pud_leaf(pud)) {
> +		spin_unlock(ptl);
> +		return 0;
> +	}
> +	stat = -ENOENT;
> +	if (pud_present(pud)) {
> +		page = vm_normal_page_pud(walk->vma, addr, pud);
> +		if (page)
> +			stat = folio_stat(page_folio(page));
> +	}
> +	pages_stat_set(walk->private, addr, end, stat);
> +	spin_unlock(ptl);
> +	walk->action = ACTION_CONTINUE;
> +	return 0;
> +}
> +
> +static int pages_stat_pmd_entry(pmd_t *pmdp, unsigned long addr,
> +				unsigned long end, struct mm_walk *walk)
> +{
> +	struct vm_area_struct *vma = walk->vma;
> +	struct page *page;
> +	spinlock_t *ptl;
> +	pte_t *ptep;
> +	pmd_t pmd;
> +	int stat;
> +
> +	pmd = pmdp_get_lockless(pmdp);
> +	if (IS_ENABLED(CONFIG_PGTABLE_HAS_HUGE_LEAVES) &&
> +	    (!pmd_present(pmd) || pmd_leaf(pmd))) {
> +		ptl = pmd_lock(walk->mm, pmdp);
> +		pmd = pmdp_get(pmdp);
> +		if (pmd_present(pmd) && !pmd_leaf(pmd)) {
> +			spin_unlock(ptl);
> +			goto pte_table;
> +		}
> +		stat = -ENOENT;
> +		if (pmd_present(pmd)) {
> +			page = vm_normal_page_pmd(vma, addr, pmd);
> +			if (page)
> +				stat = folio_stat(page_folio(page));
> +			else if (is_huge_zero_pmd(pmd))
> +				stat = -EFAULT;
> +		}
> +		pages_stat_set(walk->private, addr, end, stat);
> +		spin_unlock(ptl);
> +		return 0;
> +	}
> +
> +pte_table:
> +	ptep = pte_offset_map_lock(walk->mm, pmdp, addr, &ptl);
> +	if (!ptep) {
> +		walk->action = ACTION_AGAIN;
> +		return 0;
> +	}
> +	for (; addr < end; addr += PAGE_SIZE, ptep++) {
> +		pte_t pte = ptep_get(ptep);
> +
> +		stat = -ENOENT;
> +		if (pte_present(pte)) {
> +			page = vm_normal_page(vma, addr, pte);
> +			if (page)
> +				stat = folio_stat(page_folio(page));
> +			else if (is_zero_pfn(pte_pfn(pte)))
> +				stat = -EFAULT;
> +		}
> +		pages_stat_set(walk->private, addr, addr + PAGE_SIZE, stat);
> +	}
> +	pte_unmap_unlock(ptep - 1, ptl);
> +	return 0;
> +}
> +
> +static int pages_stat_hugetlb_entry(pte_t *ptep, unsigned long hmask,
> +				    unsigned long addr, unsigned long end,
> +				    struct mm_walk *walk)
> +{
> +#ifdef CONFIG_HUGETLB_PAGE
> +	spinlock_t *ptl;
> +	pte_t pte;
> +	int stat = -ENOENT;
> +
> +	ptl = huge_pte_lock(hstate_vma(walk->vma), walk->mm, ptep);
> +	pte = huge_ptep_get(walk->mm, addr, ptep);
> +	if (pte_present(pte))
> +		stat = folio_stat(pfn_folio(pte_pfn(pte)));
> +	pages_stat_set(walk->private, addr, end, stat);
> +	spin_unlock(ptl);
> +#endif
> +	return 0;
> +}
> +
> +static int pages_stat_pte_hole(unsigned long addr, unsigned long end,
> +			       int depth, struct mm_walk *walk)
> +{
> +	/* No VMA at all is -EFAULT, a VMA without the page is -ENOENT */
> +	pages_stat_set(walk->private, addr, end, walk->vma ? -ENOENT : -EFAULT);
> +	return 0;
> +}
> +
> +static const struct mm_walk_ops pages_stat_walk_ops = {
> +	.pud_entry	= pages_stat_pud_entry,
> +	.pmd_entry	= pages_stat_pmd_entry,
> +	.hugetlb_entry	= pages_stat_hugetlb_entry,
> +	.pte_hole	= pages_stat_pte_hole,
> +	.walk_lock	= PGWALK_RDLOCK,
> +};
> +
>  /*
>   * Determine the nodes of an array of pages and store it in an array of status.
>   */
>  static void do_pages_stat_array(struct mm_struct *mm, unsigned long nr_pages,
>  				const void __user **pages, int *status)
>  {
> -	unsigned long i;
> +	unsigned long i, n;
>  
>  	mmap_read_lock(mm);
>  
> -	for (i = 0; i < nr_pages; i++) {
> -		unsigned long addr = (unsigned long)(*pages);
> -		struct vm_area_struct *vma;
> -		struct folio_walk fw;
> -		struct folio *folio;
> -		int err = -EFAULT;
> +	for (i = 0; i < nr_pages; i += n) {
> +		unsigned long addr = (unsigned long)pages[i] & PAGE_MASK;
> +		struct pages_stat_walk psw = {
> +			.start = addr,
> +			.status = status + i,
> +		};
>  
> -		vma = vma_lookup(mm, addr);
> -		if (!vma)
> -			goto set_status;
> +		/* Walk runs of consecutive pages in one go */
> +		for (n = 1; i + n < nr_pages; n++) {
> +			unsigned long next = (unsigned long)pages[i + n] & PAGE_MASK;
>  
> -		folio = folio_walk_start(&fw, vma, addr, FW_ZEROPAGE);
> -		if (folio) {
> -			if (is_zero_folio(folio) || is_huge_zero_folio(folio))
> -				err = -EFAULT;
> -			else if (folio_is_zone_device(folio))
> -				err = -ENOENT;
> -			else
> -				err = folio_nid(folio);
> -			folio_walk_end(&fw, vma);
> -		} else {
> -			err = -ENOENT;
> +			if (next != addr + n * PAGE_SIZE || next < addr)
> +				break;
>  		}
> -set_status:
> -		*status = err;
> -
> -		pages++;
> -		status++;
> +		if (walk_page_range(mm, addr, addr + n * PAGE_SIZE,
> +				    &pages_stat_walk_ops, &psw))
> +			pages_stat_set(&psw, addr, addr + n * PAGE_SIZE, -EFAULT);
>  	}
>  
>  	mmap_read_unlock(mm);

  parent reply	other threads:[~2026-10-05 13:24 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-02  1:25 [PATCH v2 0/2] mm/migrate: speed up move_pages() node queries Qiliang Yuan
2026-10-02  1:25 ` [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array() Qiliang Yuan
2026-10-02 18:53   ` David Hildenbrand (Arm)
2026-10-05 13:24   ` Huang, Ying [this message]
2026-10-02  1:25 ` [PATCH v2 2/2] mm/migrate: raise the do_pages_stat() chunk to 512 pages Qiliang Yuan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=87v77ggu14.fsf@DESKTOP-5N7EMDA \
    --to=ying.huang@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=byungchul@sk.com \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=joshua.hahnjy@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=matthew.brost@intel.com \
    --cc=odys.yuan@gmail.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®