Bug 277476 - graphics/drm-515-kmod: amdgpu periodic hangs due to phys contig allocations
Summary: graphics/drm-515-kmod: amdgpu periodic hangs due to phys contig allocations
Status: Closed FIXED
Alias: None
Product: Base System
Classification: Unclassified
Component: kern (show other bugs)
Version: CURRENT
Hardware: Any Any
: --- Affects Only Me
Assignee: Olivier Certner
URL: https://github.com/freebsd/drm-kmod/i...
Keywords:
Depends on:
Blocks:
 
Reported: 2024-03-04 14:12 UTC by Josef 'Jeff' Sipek
Modified: 2025-10-03 19:58 UTC (History)
14 users (show)

See Also:
linimon: maintainer-feedback? (x11)
olce: mfc-stable15+
olce: mfc-stable14+


Attachments
PR277476 fix (2.87 KB, patch)
2024-11-14 00:43 UTC, sigsys
no flags Details | Diff
dtrace profile (215.94 KB, text/plain)
2025-03-12 20:09 UTC, Ivan Rozhuk
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description Josef 'Jeff' Sipek 2024-03-04 14:12:20 UTC
Two weeks ago I replaced an ancient nvidia graphics card with an AMD RX580 card to run open source drivers. Everything works fine most of the time, but occasionally the system hangs for a few seconds (5-10, usually).  The longer the system has been up, the worse it gets.

Digging into it a bit, it is because userspace (it always looks like X) does an ioctl into drm which then tries to allocate a large-ish piece of physically contiguous memory.  This explains why it gets worse as uptime increases (free physical memory fragmentation) and when running firefox (the most memory hungry application I use).  I know nothing about graphics cards, the software stack supporting them, or the linux kernel API compatibility layer, but clearly it'd be beneficial if amdgpu/drm/whatever could make use of *virtually* contiguous pages or some kind of allocation caching/reuse to avoid repeatedly asking the vm code for physically contiguous ranges.

To conclude the above, I did a handful of dtrace-based experiments.

While one of the "temporary hangs" was happening, the following was the most common (non-idle) profiler stack:

# dtrace -n 'profile-97{@[stack()]=count()}'
...
              kernel`vm_phys_alloc_contig+0x11d
              kernel`linux_alloc_pages+0x8f
              ttm.ko`ttm_pool_alloc+0x2cb
              ttm.ko`ttm_tt_populate+0xc5
              ttm.ko`ttm_bo_handle_move_mem+0xc3
              ttm.ko`ttm_bo_validate+0xb4
              ttm.ko`ttm_bo_init_reserved+0x199
              amdgpu.ko`amdgpu_bo_create+0x1eb
              amdgpu.ko`amdgpu_bo_create_user+0x21
              amdgpu.ko`amdgpu_gem_create_ioctl+0x1e2
              drm.ko`drm_ioctl_kernel+0xc6
              drm.ko`drm_ioctl+0x2b5
              kernel`linux_file_ioctl+0x312
              kernel`kern_ioctl+0x255
              kernel`sys_ioctl+0x123
              kernel`amd64_syscall+0x109
              kernel`0xffffffff80fe43eb

The latency of vm_phys_alloc_contig (entry to return) is bimodal - with latencies in the single digit *milli*seconds during the "temporary hangs":

# dtrace -n 'fbt::vm_phys_alloc_contig:entry{self->ts=timestamp}' -n 'fbt::vm_phys_alloc_contig:return/self->ts/{this->delta=timestamp-self->ts; @=quantize(this->delta);}' -n 'tick-1sec{printa(@)}'
...
           value  ------------- Distribution ------------- count    
             256 |                                         0        
             512 |@                                        2606     
            1024 |@@@@@@@@@                                18207    
            2048 |@                                        2534     
            4096 |                                         894      
            8192 |                                         34       
           16384 |                                         78       
           32768 |                                         58       
           65536 |                                         219      
          131072 |                                         306      
          262144 |                                         310      
          524288 |                                         735      
         1048576 |                                         174      
         2097152 |@@                                       4364     
         4194304 |@@@@@@@@@@@@@@@@@@@@@@@@                 47475    
         8388608 |@                                        1546     
        16777216 |                                         2        
        33554432 |                                         0     

The number of pages being allocated:

# dtrace -n 'fbt::vm_phys_alloc_contig:entry/arg1>1/{@=quantize(arg1)}' -n 'tick-1sec{printa(@)}'
...
           value  ------------- Distribution ------------- count    
               1 |                                         0        
               2 |@@@                                      15       
               4 |@                                        7        
               8 |@@@                                      16       
              16 |@@                                       10       
              32 |@@                                       10       
              64 |@                                        7        
             128 |@@@                                      12       
             256 |@@@                                      12       
             512 |@@@@@@@@@@@@@@                           68       
            1024 |@@@@@@@                                  32       
            2048 |                                         0      

I did a few more dtrace experiments, but they all point to the same thing - a drm/amdgpu related ioctl wants 4MB of physically contiguous memory often enough to become a headache.  4MB isn't too much given than the system has 32GB of RAM, but physically contiguous takes a while to fulfill sometimes.


The card:

vgapci0@pci0:1:0:0:     class=0x030000 rev=0xe7 hdr=0x00 vendor=0x1002 device=0x67df subvendor=0x1da2 subdevice=0xe353
    vendor     = 'Advanced Micro Devices, Inc. [AMD/ATI]'
    device     = 'Ellesmere [Radeon RX 470/480/570/570X/580/580X/590]'
    class      = display
    subclass   = VGA

$ pkg info|grep -i amd             
gpu-firmware-amd-kmod-aldebaran-20230625 Firmware modules for aldebaran AMD GPUs
gpu-firmware-amd-kmod-arcturus-20230625 Firmware modules for arcturus AMD GPUs
gpu-firmware-amd-kmod-banks-20230625 Firmware modules for banks AMD GPUs
gpu-firmware-amd-kmod-beige-goby-20230625 Firmware modules for beige_goby AMD GPUs
gpu-firmware-amd-kmod-bonaire-20230625 Firmware modules for bonaire AMD GPUs
gpu-firmware-amd-kmod-carrizo-20230625 Firmware modules for carrizo AMD GPUs
gpu-firmware-amd-kmod-cyan-skillfish2-20230625 Firmware modules for cyan_skillfish2 AMD GPUs
gpu-firmware-amd-kmod-dimgrey-cavefish-20230625 Firmware modules for dimgrey_cavefish AMD GPUs
gpu-firmware-amd-kmod-fiji-20230625 Firmware modules for fiji AMD GPUs
gpu-firmware-amd-kmod-green-sardine-20230625 Firmware modules for green_sardine AMD GPUs
gpu-firmware-amd-kmod-hainan-20230625 Firmware modules for hainan AMD GPUs
gpu-firmware-amd-kmod-hawaii-20230625 Firmware modules for hawaii AMD GPUs
gpu-firmware-amd-kmod-kabini-20230625 Firmware modules for kabini AMD GPUs
gpu-firmware-amd-kmod-kaveri-20230625 Firmware modules for kaveri AMD GPUs
gpu-firmware-amd-kmod-mullins-20230625 Firmware modules for mullins AMD GPUs
gpu-firmware-amd-kmod-navi10-20230625 Firmware modules for navi10 AMD GPUs
gpu-firmware-amd-kmod-navi12-20230625 Firmware modules for navi12 AMD GPUs
gpu-firmware-amd-kmod-navi14-20230625 Firmware modules for navi14 AMD GPUs
gpu-firmware-amd-kmod-navy-flounder-20230625 Firmware modules for navy_flounder AMD GPUs
gpu-firmware-amd-kmod-oland-20230625 Firmware modules for oland AMD GPUs
gpu-firmware-amd-kmod-picasso-20230625 Firmware modules for picasso AMD GPUs
gpu-firmware-amd-kmod-pitcairn-20230625 Firmware modules for pitcairn AMD GPUs
gpu-firmware-amd-kmod-polaris10-20230625 Firmware modules for polaris10 AMD GPUs
gpu-firmware-amd-kmod-polaris11-20230625 Firmware modules for polaris11 AMD GPUs
gpu-firmware-amd-kmod-polaris12-20230625 Firmware modules for polaris12 AMD GPUs
gpu-firmware-amd-kmod-raven-20230625 Firmware modules for raven AMD GPUs
gpu-firmware-amd-kmod-raven2-20230625 Firmware modules for raven2 AMD GPUs
gpu-firmware-amd-kmod-renoir-20230625 Firmware modules for renoir AMD GPUs
gpu-firmware-amd-kmod-si58-20230625 Firmware modules for si58 AMD GPUs
gpu-firmware-amd-kmod-sienna-cichlid-20230625 Firmware modules for sienna_cichlid AMD GPUs
gpu-firmware-amd-kmod-stoney-20230625 Firmware modules for stoney AMD GPUs
gpu-firmware-amd-kmod-tahiti-20230625 Firmware modules for tahiti AMD GPUs
gpu-firmware-amd-kmod-tonga-20230625 Firmware modules for tonga AMD GPUs
gpu-firmware-amd-kmod-topaz-20230625 Firmware modules for topaz AMD GPUs
gpu-firmware-amd-kmod-vangogh-20230625 Firmware modules for vangogh AMD GPUs
gpu-firmware-amd-kmod-vega10-20230625 Firmware modules for vega10 AMD GPUs
gpu-firmware-amd-kmod-vega12-20230625 Firmware modules for vega12 AMD GPUs
gpu-firmware-amd-kmod-vega20-20230625 Firmware modules for vega20 AMD GPUs
gpu-firmware-amd-kmod-vegam-20230625 Firmware modules for vegam AMD GPUs
gpu-firmware-amd-kmod-verde-20230625 Firmware modules for verde AMD GPUs
gpu-firmware-amd-kmod-yellow-carp-20230625 Firmware modules for yellow_carp AMD GPUs
suitesparse-amd-3.3.0          Symmetric approximate minimum degree
suitesparse-camd-3.3.0         Symmetric approximate minimum degree
suitesparse-ccolamd-3.3.0      Constrained column approximate minimum degree ordering
suitesparse-colamd-3.3.0       Column approximate minimum degree ordering algorithm
webcamd-5.17.1.2_1             Port of Linux USB webcam and DVB drivers into userspace
xf86-video-amdgpu-22.0.0_1     X.Org amdgpu display driver
$ pkg info|grep -i drm
drm-515-kmod-5.15.118_3        DRM drivers modules
drm-kmod-20220907_1            Metaport of DRM modules for the linuxkpi-based KMS components
gpu-firmware-kmod-20230210_1,1 Firmware modules for the drm-kmod drivers
libdrm-2.4.120_1,1             Direct Rendering Manager library and headers
Comment 1 Josef 'Jeff' Sipek 2024-03-04 14:28:43 UTC
While gathering all the dtrace data, I was so distracted I forgot to mention:

$ freebsd-version -kru
14.0-RELEASE-p5
14.0-RELEASE-p5
14.0-RELEASE-p5
Comment 2 Josef 'Jeff' Sipek 2024-04-06 14:22:49 UTC
I dug a bit more into this.  It looks like the drm code has provisions for
allocating memory via dma APIs.  The FreeBSD port doesn't implement those.

Specifically, looking at drm-kmod-drm_v5.15.25_5 source:

drivers/gpu/drm/amd/amdgpu/gmc_v*.c sets adev->need_swiotlb to
drm_need_swiotlb(...).  drm_need_swiotlb is implemented in
drivers/gpu/drm/drm_cache.c as a 'return false' on FreeBSD.

Later on, amdgpu_ttm_init calls ttm_device_init with the use_dma_alloc
argument equal to adev->need_swiotlb (IOW, false).

Much later on, ttm_pool_alloc is called to allocate a buffer.  That in turn
calls ttm_pool_alloc_page which amounts to:

	if (!use_dma_alloc)
		return alloc_pages(...);
	
	panic("ttm_pool.c: use_dma_alloc not implemented");

So, because of the 'return false' during initialization, we always call
alloc_pages (aka. linux_alloc_pages) which tries to allocate physically
contiguous memory.

As I said before, I don't know anything about the graphics stack, so it is
possible that this dma API is completely irrelevant.


Looking at ttm_pool_alloc some more, it immediately turns the physically
contiguous allocation into an array of struct page pointers (tt->pagse).
So, depending on how the rest of the module uses the buffer & pages, it
may be relatively easy to switch to a virtually-contiguous allocation.
Comment 3 Tomasz "CeDeROM" CEDRO 2024-05-15 23:23:56 UTC
I also have RX580. After upgrading from 13.2 to 14.0 I got frequent kernel panics. 

Related reports:
* https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=276985
* https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=278212

I noticed that setting manual xorg.conf admgpu DRI=2 make kernel panic less frequent but instead system got slower and slower until unresponsive, if I managed to kill xorg in time it could work for a while again until I had to kill xorg again.

I could not work out fallback with safe non accelerated xorg.conf using scfb that would allow dual screen setup with one screen rotated. Dual monitor is only possible with amdgpu loaded. It was also not possible to disable acceleration in amdgpu and have xrandr (secondary screen rotation).

Rolled back to 13.2. DRM 5.15 / AMDGPU / LinuxKPI makes 14.0 unreliable.
Comment 4 Olivier Certner freebsd_committer freebsd_triage 2024-09-02 19:04:09 UTC
(In reply to Josef 'Jeff' Sipek from comment #0)

I reported this independently in the drm-kmod GitHub project as https://github.com/freebsd/drm-kmod/issues/302.  Going to also cross-link this PR there.

I'd really like to get to the bottom of this.  However, I don't plan to have the time to do so before end of month at the very least.

drm-61-kmod exhibits the same problem.  However, drm-510-kmod works fine for me.

(In reply to Tomasz "CeDeROM" CEDRO from comment #3)

Please see my comment 8 in bug #278212.
Comment 5 sigsys 2024-11-08 09:04:51 UTC
Yeah so this problem was super annoying. But thanks to the information already posted here, seems like it wasn't too hard to fix.

IIUC the drm code (ttm_pool_alloc()) asking for contiguous pages doesn't actually need contiguous pages. It's just an opportunistic optimization. When allocation fails, it fallsback to asking for less and less contiguous pages (eventually only asking for one page at a time). When ttm_pool_alloc_page() asks for more than one page, it passes alloc_pages() some extra flags (__GFP_NOMEMALLOC | __GFP_NORETRY | __GFP_NOWARN | __GFP_KSWAPD_RECLAIM).

What's expensive is the vm_page_reclaim_contig() in linux_alloc_pages(). The function tries too hard to find contiguous memory (that the drm code doesn't even require) and as physical memory gets too fragmented it becomes very slow.

So, very simple fix, make linux_alloc_pages() react to one of the flag passed by the drm code:

diff --git a/sys/compat/linuxkpi/common/include/linux/gfp.h b/sys/compat/linuxkpi/common/include/linux/gfp.h
index 2fcc0dc05f29..58a021086c98 100644
--- a/sys/compat/linuxkpi/common/include/linux/gfp.h
+++ b/sys/compat/linuxkpi/common/include/linux/gfp.h
@@ -44,7 +44,6 @@
 #define	__GFP_NOWARN	0
 #define	__GFP_HIGHMEM	0
 #define	__GFP_ZERO	M_ZERO
-#define	__GFP_NORETRY	0
 #define	__GFP_NOMEMALLOC 0
 #define	__GFP_RECLAIM   0
 #define	__GFP_RECLAIMABLE   0
@@ -58,7 +57,8 @@
 #define	__GFP_KSWAPD_RECLAIM	0
 #define	__GFP_WAIT	M_WAITOK
 #define	__GFP_DMA32	(1U << 24) /* LinuxKPI only */
-#define	__GFP_BITS_SHIFT 25
+#define	__GFP_NORETRY	(1U << 25) /* LinuxKPI only */
+#define	__GFP_BITS_SHIFT 26
 #define	__GFP_BITS_MASK	((1 << __GFP_BITS_SHIFT) - 1)
 #define	__GFP_NOFAIL	M_WAITOK
 
diff --git a/sys/compat/linuxkpi/common/src/linux_page.c b/sys/compat/linuxkpi/common/src/linux_page.c
index 18b90b5e3d73..71a6890a3795 100644
--- a/sys/compat/linuxkpi/common/src/linux_page.c
+++ b/sys/compat/linuxkpi/common/src/linux_page.c
@@ -118,7 +118,7 @@ linux_alloc_pages(gfp_t flags, unsigned int order)
 			page = vm_page_alloc_noobj_contig(req, npages, 0, pmax,
 			    PAGE_SIZE, 0, VM_MEMATTR_DEFAULT);
 			if (page == NULL) {
-				if (flags & M_WAITOK) {
+				if ((flags & (M_WAITOK | __GFP_NORETRY)) == M_WAITOK) {
 					int err = vm_page_reclaim_contig(req,
 					    npages, 0, pmax, PAGE_SIZE, 0);
 					if (err == ENOMEM)

Been working fine here with amdgpu for about 3 weeks.

(The drm modules need to be recompiled with the modified kernel header.)
Comment 6 Emmanuel Vadot freebsd_committer freebsd_triage 2024-11-10 11:00:44 UTC
Intersting find, I've never could reproduce this bug on my RX550, Olivier can you test if the code (which looks ok to me) fixes the issue for you ?
Comment 7 rk 2024-11-10 19:55:42 UTC
(In reply to sigsys from comment #5)

I've been suffering from this issue on a Ryzen 9 4900H with embedded
Renoir graphics using drm-61-kmod-6.1.92_2 on stable/14-n268738-048132192698.
Playing videos using mpv easily triggered the slowdown after some time
(especially 4K videos).
I've implemented the suggested fix and now I cannot reproduce the behaviour
anymore (tested for 2 days now). Even playing multiple 4K videos in
parallel does not cause the problem.
Thanks for the fix.
Comment 8 Olivier Certner freebsd_committer freebsd_triage 2024-11-12 09:43:41 UTC
(In reply to sigsys from comment #5)

> IIUC the drm code (ttm_pool_alloc()) asking for contiguous pages doesn't actually need contiguous pages. It's just an opportunistic optimization.

That would be very good news (at least from the users' point of view).

Have not spent time on this issue since my last posts.  I had naively thought that the new DRM ports really needed contiguous allocation for whatever reason, and should probably have looked a bit further instead of assuming this would need some deep and highly time consuming analysis.

(In reply to Emmanuel Vadot from comment #6)

Will test that soon and report.
Comment 9 Emmanuel Vadot freebsd_committer freebsd_triage 2024-11-12 12:23:19 UTC
(In reply to sigsys from comment #5)

Waiting for more people to test but in the meantime could you add a git-format patch to this bug please ? (So with full commit message and correct authorship).
Comment 10 Pierre Pronchery 2024-11-13 19:05:38 UTC
(In reply to Emmanuel Vadot from comment #6)
The patch also works well for me, no slowdowns to report after 24 hours.
Comment 11 Eirik Oeverby 2024-11-13 20:49:14 UTC
(In reply to sigsys from comment #5)
Has this patch landed already? I'm eager to test on my threadripper with Navi 24 [Radeon PRO W6400]; it's borderline useless after ~48h uptime and needs frequent reboots to fix. At least it's better than before I clamped the ARC to 8GB to slow the process down..
Comment 12 sigsys 2024-11-14 00:43:11 UTC
Created attachment 255155 [details]
PR277476 fix
Comment 13 sigsys 2024-11-14 00:48:43 UTC
(In reply to Emmanuel Vadot from comment #9)
Alright here it is.

Is it already too late to have this merged in 14.2?

I'm pretty sure this patch is safe. GFP_NORETRY isn't used in-tree at all right now. And this patch makes it do pretty much what it says. It doesn't retry. You'd hope that any code using this flag would expect allocations to fail...

The problem doesn't always happen for everyone but when it does man it's rough. After a week or two I was getting hangs that lasted 15 seconds sometimes. Restarting firefox would fix it for a while but eventually it becomes unusable.

Even if this made it in 14.2 IIUC it would take a while before the 14.X packages would be compiled against the new kernel headers, but it would already be useful to have it in base so that you could get the fix by compiling drm-kmod from ports.
Comment 14 Matthew D. Fuller 2024-11-14 02:08:03 UTC
FWIW, I definitely ran into what sounds just like this (with several different cards, on both 515 and 61; 510 was always rock solid).  After a few days, I'd sometimes get freezes lasting a minute or more.

A workaround that seems to work for me has been switching from the amdgpu to the modesetting X driver; I still occasionally see little blips, but they resolve and don't seem to pile up the way they did on amdgpu, even after months of uptime.
Comment 15 Olivier Certner freebsd_committer freebsd_triage 2024-11-14 09:24:12 UTC
(In reply to sigsys from comment #13)

It seems manu@ is having a crash with the patch applied on 5.15.  So while it seems safe, we have to rule out some possible impacts in certain situations.

I'm afraid it is too late to have it merged in 14.2 anyway, so let's be sure we are not regressing anything while fixing the problem.
Comment 16 Ivan Rozhuk 2025-03-09 21:03:17 UTC
Never see this issue on "AMD Ryzen 5 2500U with Radeon Vega Mobile Gfx", but always on
5950x + RX 5600 XT.
Software config identical, FBSD 14/stable.
xorg + amdgpu x driver.


To reduce freezes I use:
- picom
- cpuset to leave 8 cores free while build ports
- script that creates 16gb file on tmpfs and remove it

Script makes free ~20gb and almost no freezes until freemem > 3gb.
But after ~1 month of uptime I start see freezes even with freemem > 20gb.
Even without debug tools this was looks like memory fragmentation issue. :)


Is it possible to implement some memory defrag code in pagedaemon?
vm_page_reclaim_contig() also used by iommu, ktls and shm so improving it will make FBSD better even for server roles.



(In reply to Josef 'Jeff' Sipek from comment #0)

Thanks for debugging!


(In reply to sigsys from comment #5)

Thanks for patch, I will test it, but it requires at least 2 weeks to make sure that freezes go away.



(In reply to Emmanuel Vadot from comment #6)

Do you use xorg + amdgpu x driver?
Comment 17 Olivier Certner freebsd_committer freebsd_triage 2025-03-10 10:33:28 UTC
(In reply to Ivan Rozhuk from comment #16)

> Thanks for patch, I will test it, but it requires at least 2 weeks to make sure that freezes go away.

Yes, and please report about your experience.

I'll do an extra test on my side.

Unless something goes wrong, I'd like to move this forward soon, and foremost before we start the release process for 14.3.
Comment 18 Evgenii Khramtsov 2025-03-10 13:45:22 UTC
(In reply to Olivier Certner from comment #17)

> Yes, and please report about your experience.

My 2 cents: I have used the patch for months with 6.1 back when it was tip of drm-kmod, then with graphics/drm-66-kmod, and the patch doesn't panic my desktop, and resolves the issue for me. I have never used it with anything less than 6.1.
Comment 19 Matthew D. Fuller 2025-03-10 15:11:04 UTC
I've run with the patch (slightly massaged to fit stable/14) with amdgpu and 6.1 for a month without any hint of the freezes showing up, so it certainly feels like a fix here.
Comment 20 Ivan Rozhuk 2025-03-12 20:09:54 UTC
Created attachment 258607 [details]
dtrace profile

Patch did not help, at least in case: xorg + amdgpu xdriver.
This how it was landed on 14/stable: https://github.com/rozhuk-im/freebsd/commit/b739c10c50aa37e247dc95f7b93f6fe58d86016d


I have attached dtrace profile output that captured while freezes happen.
I do not see here vm_phys_alloc_contig() after ttm_pool_alloc(), probably -O2/-O3 opt level "optimize" out it.

Here few new things that show increased latency on freezes:
(I do not collect many freezes, in some tests only few freezes collected)

              kernel`lock_delay+0x12
              amdgpu.ko`amdgpu_gem_fault+0x86
              kernel`linux_cdev_pager_populate+0x128
              kernel`vm_fault_allocate+0x185
              kernel`vm_fault+0x39c
              kernel`vm_fault_trap+0x4c
              kernel`trap_pfault+0x20a
              kernel`trap+0x4a8
              kernel`0xffffffff80a11ca8
               20
dtrace -n 'fbt::amdgpu_gem_fault:entry{self->ts=timestamp}' -n 'fbt::amdgpu_gem_fault:return/self->ts/{this->delta=timestamp-self->ts; @=quantize(this->delta);}' -n 'tick-1sec{printa(@)}'
  0  66064                       :tick-1sec 

           value  ------------- Distribution ------------- count    
             512 |                                         0        
            1024 |                                         1        
            2048 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@       1357     
            4096 |@@@                                      110      
            8192 |@@@                                      103      
           16384 |                                         4        
           32768 |                                         3        
           65536 |                                         0        
          131072 |                                         0        
          262144 |                                         0        
          524288 |                                         0        
         1048576 |                                         0        
         2097152 |                                         1        
         4194304 |                                         3        
         8388608 |                                         2        
        16777216 |                                         0        
        33554432 |                                         0        
        67108864 |                                         0        
       134217728 |                                         0        
       268435456 |                                         1        
       536870912 |                                         1        
      1073741824 |                                         1        
      2147483648 |                                         0        



              kernel`lock_delay+0x14
              kernel`malloc_large+0x2c
              kernel`lkpi_kmalloc_cb+0x44
              kernel`lkpi_kmalloc+0x27
              amdgpu.ko`dc_create_state+0x18
              amdgpu.ko`amdgpu_dm_atomic_commit_tail+0xd4
              drm.ko`commit_tail+0xa7
              kernel`linux_work_fn+0xed
              kernel`taskqueue_run_locked+0x187
              kernel`taskqueue_thread_loop+0xc2
              kernel`fork_exit+0x86
              kernel`0xffffffff80a12d0e
               88
dtrace -n 'fbt::dc_create_state:entry{self->ts=timestamp}' -n 'fbt::dc_create_state:return/self->ts/{this->delta=timestamp-self->ts; @=quantize(this->delta);}' -n 'tick-1sec{printa(@)}'
  0  66064                       :tick-1sec 

           value  ------------- Distribution ------------- count    
            4096 |                                         0        
            8192 |                                         2        
           16384 |@@@@@@@@@@@@@@@@@@@@@                    1271     
           32768 |@@@@@@@@@@@@@@@@@@                       1087     
           65536 |                                         30       
          131072 |                                         1        
          262144 |                                         0        
          524288 |                                         3        
         1048576 |                                         4        
         2097152 |                                         2        
         4194304 |                                         0        
         8388608 |                                         0        
        16777216 |                                         0        
        33554432 |                                         0        
        67108864 |                                         0        
       134217728 |                                         1        
       268435456 |                                         1        
       536870912 |                                         5        
      1073741824 |                                         4        
      2147483648 |                                         0        


              kernel`lock_delay+0x14
              kernel`free+0x9b
              amdgpu.ko`amdgpu_dm_atomic_commit_tail+0x2f9a
              drm.ko`commit_tail+0xa7
              kernel`linux_work_fn+0xed
              kernel`taskqueue_run_locked+0x187
              kernel`taskqueue_thread_loop+0xc2
              kernel`fork_exit+0x86
              kernel`0xffffffff809aaf6e
              399
dtrace -n 'fbt::amdgpu_dm_atomic_commit_tail:entry{self->ts=timestamp}' -n 'fbt::amdgpu_dm_atomic_commit_tail:return/self->ts/{this->delta=timestamp-self->ts; @=quantize(this->delta);}' -n 'tick-1sec{printa(@)}'
  0  66190                       :tick-1sec 

           value  ------------- Distribution ------------- count    
           16384 |                                         0        
           32768 |                                         4        
           65536 |                                         6        
          131072 |                                         2        
          262144 |                                         0        
          524288 |                                         6        
         1048576 |                                         15       
         2097152 |@                                        29       
         4194304 |@@@                                      106      
         8388608 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@       1323     
        16777216 |@                                        44       
        33554432 |                                         1        
        67108864 |                                         2        
       134217728 |                                         4        
       268435456 |                                         8        
       536870912 |                                         5        
      1073741824 |                                         5        
      2147483648 |                                         0        


              kernel`lock_delay+0x14
              kernel`zone_import+0xf2
              kernel`cache_alloc+0x309
              kernel`cache_alloc_retry+0x2c
              kernel`malloc+0x48
              ttm.ko`ttm_sg_tt_init+0x61
              amdgpu.ko`amdgpu_ttm_tt_create+0x4a
              ttm.ko`ttm_tt_create+0x4e
              ttm.ko`ttm_bo_validate+0x60
              ttm.ko`ttm_bo_init_reserved+0x194
              amdgpu.ko`amdgpu_bo_create+0x295
              amdgpu.ko`amdgpu_bo_create_user+0x21
              amdgpu.ko`amdgpu_gem_userptr_ioctl+0x82
              drm.ko`drm_ioctl_kernel+0xbc
              drm.ko`drm_ioctl+0x25e
              kernel`linux_file_ioctl+0x30f
              kernel`kern_ioctl+0x1b0
              kernel`sys_ioctl+0x117
              kernel`amd64_syscall+0xeb
              kernel`0xffffffff809aa81b
               46
dtrace -n 'fbt::amdgpu_ttm_tt_create:entry{self->ts=timestamp}' -n 'fbt::amdgpu_ttm_tt_create:return/self->ts/{this->delta=timestamp-self->ts; @=quantize(this->delta);}' -n 'tick-1sec{printa(@)}'
  0  66190                       :tick-1sec 

           value  ------------- Distribution ------------- count    
             128 |                                         0        
             256 |                                         4        
             512 |@@@@@@@@@@                               5764     
            1024 |@@@@@@@@@@@@                             6635     
            2048 |@@@@@@@@@                                5087     
            4096 |@@@@@@@@                                 4334     
            8192 |@@                                       875      
           16384 |                                         72       
           32768 |                                         9        
           65536 |                                         3        
          131072 |                                         0        
(this looks ok)


dtrace -n 'fbt::amdgpu_bo_create:entry{self->ts=timestamp}' -n 'fbt::amdgpu_bo_create:return/self->ts/{this->delta=timestamp-self->ts; @=quantize(this->delta);}' -n 'tick-1sec{printa(@)}'
  0  66190                       :tick-1sec 

           value  ------------- Distribution ------------- count    
             256 |                                         0        
             512 |                                         2        
            1024 |@@@@@@                                   2303     
            2048 |@@@@@@@@@@                               4190     
            4096 |@@@@@@@@@@@@                             4800     
            8192 |@@@@@@@@@@                               4002     
           16384 |@@                                       845      
           32768 |                                         124      
           65536 |                                         39       
          131072 |                                         20       
          262144 |                                         4        
          524288 |                                         9        
         1048576 |                                         2        
         2097152 |                                         3        
         4194304 |                                         8        
         8388608 |                                         5        
        16777216 |                                         0        
        33554432 |                                         1        
        67108864 |                                         2        
       134217728 |                                         1        
       268435456 |                                         0        
       536870912 |                                         2        
      1073741824 |                                         0        
      2147483648 |                                         1        
      4294967296 |                                         0        

dtrace -n 'fbt::add_hole:entry{self->ts=timestamp}' -n 'fbt::add_hole:return/self->ts/{this->delta=timestamp-self->ts; @=quantize(this->delta);}' -n 'tick-1sec{printa(@)}'
  0  66190                       :tick-1sec 

           value  ------------- Distribution ------------- count    
             128 |                                         0        
             256 |@@@@@@@@@@@@                             5762     
             512 |@@@@@@@@@@@@@                            6548     
            1024 |@@@@@@@                                  3287     
            2048 |@@@@@                                    2648     
            4096 |@@@                                      1508     
            8192 |                                         105      
           16384 |                                         11       
           32768 |                                         4        
           65536 |                                         1        
          131072 |                                         0        
(this looks ok)

dtrace -n 'fbt::ttm_pool_alloc:entry{self->ts=timestamp}' -n 'fbt::ttm_pool_alloc:return/self->ts/{this->delta=timestamp-self->ts; @=quantize(this->delta);}' -n 'tick-1sec{printa(@)}'
  0  66190                       :tick-1sec 

           value  ------------- Distribution ------------- count    
             128 |                                         0        
             256 |@@                                       29       
             512 |@@@@@@@@                                 96       
            1024 |@@@@@@                                   81       
            2048 |@@@@@                                    67       
            4096 |@@@@@@@@                                 106      
            8192 |@@@                                      33       
           16384 |                                         5        
           32768 |                                         1        
           65536 |@                                        17       
          131072 |@                                        12       
          262144 |                                         3        
          524288 |                                         6        
         1048576 |@                                        10       
         2097152 |@                                        13       
         4194304 |                                         2        
         8388608 |                                         2        
        16777216 |                                         3        
        33554432 |                                         5        
        67108864 |                                         3        
       134217728 |                                         4        
       268435456 |@                                        7        
       536870912 |                                         5        
      1073741824 |                                         1        
      2147483648 |                                         1        
      4294967296 |                                         0        


If some one have ideas - I can play more with dtrace and test other patches/settings.
Comment 21 Olivier Certner freebsd_committer freebsd_triage 2025-03-12 20:53:37 UTC
(In reply to Ivan Rozhuk from comment #20)

Very important: Once the patch has been applied, you both have to rebuild your kernel *and* the drm-kmod modules with the patched 'gfp.h' header.

Given previous analysis, it seems unlikely at this stage that the patch wouldn't fix what you're observing, but let's see.
Comment 22 Olivier Certner freebsd_committer freebsd_triage 2025-03-12 20:59:26 UTC
(In reply to Evgenii Khramtsov from comment #18)
(In reply to fullermd from comment #19)

I hear you.  I have been running stable/14 for months and drm-61-kmod, and it works like a charm here.

I have no doubt that this fixes the main problem, the only pending thing was to be sure that the change causes no new crashes, as manu@ hinted at that on -CURRENT.  I still have to try with a recent -CURRENT (this is the "extra test" I mentioned above).
Comment 23 Ivan Rozhuk 2025-03-12 21:02:43 UTC
(In reply to Olivier Certner from comment #21)

Is is done in auto mode by my build scripts.
I use PORTS_MODULES+= to make sure that kernel modu;es from ports auto rebuild+install with systems.


# ls /boot/kernel
...
-r--r--r--   1 root wheel   29K Mar 12 17:26:39 2025 dtaudit.ko
-r--r--r--   1 root wheel   19K Mar 12 17:26:39 2025 dtmalloc.ko
-r--r--r--   1 root wheel   30K Mar 12 17:26:39 2025 dtnfscl.ko
-r--r--r--   1 root wheel   27K Mar 12 17:26:39 2025 dtrace_test.ko
-r--r--r--   1 root wheel  374K Mar 12 17:26:39 2025 dtrace.ko
-r--r--r--   1 root wheel   16K Mar 12 17:26:39 2025 dtraceall.ko
...
-r--r--r--   1 root wheel   15M Mar 12 17:26:32 2025 kernel
-r--r--r--   1 root wheel   44K Mar 12 17:26:40 2025 kinst.ko
-r--r--r--   1 root wheel  206K Mar 12 17:26:39 2025 krpc.ko
-r--r--r--   1 root wheel   30K Mar 12 17:26:39 2025 ksyms.ko
-r--r--r--   1 root wheel   21K Mar 12 17:26:39 2025 libmchain.ko
-r--r--r--   1 root wheel   59K Mar 12 17:26:39 2025 lindebugfs.ko
-rw-r--r--   1 root wheel  125K Mar 12 17:26:41 2025 linker.hints
-r--r--r--   1 root wheel  165K Mar 12 17:26:39 2025 linux_common.ko
-r--r--r--   1 root wheel  449K Mar 12 17:26:39 2025 linux.ko
-r--r--r--   1 root wheel  414K Mar 12 17:26:39 2025 linux64.ko
-r--r--r--   1 root wheel   46K Mar 12 17:26:39 2025 linuxkpi_hdmi.ko
-r--r--r--   1 root wheel   57K Mar 12 17:26:39 2025 linuxkpi_video.ko
-r--r--r--   1 root wheel  335K Mar 12 17:26:39 2025 linuxkpi.ko
...


# ls /boot/modules/
...
-r--r--r--   1 root wheel  369K Mar 12 17:26:50 2025 amdgpu_raven_vcn_bin.ko
-r--r--r--   1 root wheel   10M Mar 12 17:26:44 2025 amdgpu.ko
...
-r--r--r--   1 root wheel  2.0M Mar 12 17:26:44 2025 radeonkms.ko
-r--r--r--   1 root wheel  100K Mar 12 17:26:44 2025 ttm.ko
Comment 24 Evgenii Khramtsov 2025-03-12 21:15:08 UTC
(In reply to Olivier Certner from comment #22)

> I still have to try with a recent -CURRENT (this is the "extra test" I mentioned above).

See bug 282605 to avoid unrelated crashes on main.

FWIW, mine "for months" means that I've been running main not older than a week with this patch, during all the mentioned time, e.g. my main now is as of base 717adecbbb52.
Comment 25 commit-hook freebsd_committer freebsd_triage 2025-03-25 09:19:59 UTC
A commit in branch main references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=718d1928f8748fe4429c011296f94f194d63c695

commit 718d1928f8748fe4429c011296f94f194d63c695
Author:     Mathieu <sigsys@gmail.com>
AuthorDate: 2024-11-14 00:24:02 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-03-25 08:41:44 +0000

    LinuxKPI: make linux_alloc_pages() honor __GFP_NORETRY

    This is to fix slowdowns with drm-kmod that get worse over time as
    physical memory become more fragmented (and probably also depending on
    other factors).

    Based on information posted in this bug report:
    https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=277476

    By default, linux_alloc_pages() retries failed allocations by calling
    vm_page_reclaim_contig() to attempt to free contiguous physical memory
    pages. vm_page_reclaim_contig() does not always succeed and calling it
    can be very slow even when it fails. When physical memory is very
    fragmented, vm_page_reclaim_contig() can end up being called (and
    failing) after every allocation attempt. This could cause very
    noticeable graphical desktop hangs (which could last seconds).

    The drm-kmod code in question attempts to allocate multiple contiguous
    pages at once but does not actually require them to be contiguous. It
    can fallback to doing multiple smaller allocations when larger
    allocations fail. It passes alloc_pages() the __GFP_NORETRY flag in this
    case.

    This patch makes linux_alloc_pages() fail early (without retrying) when
    this flag is passed.

    [olce: The problem this patch fixes is longer and longer GUI freezes as
    a machine's memory gets filled and becomes fragmented, when using amdgpu
    from DRM kmod 5.15 and DRM kmod 6.1 (DRM kmod 5.10 is unaffected; newer
    Linux kernel introduced an "optimization" by which a pool of pages is
    filled preferentially with contiguous pages, which triggered the problem
    for us).  The original commit message above evokes freezes lasting
    seconds, but I occasionally witnessed some lasting tens of minutes,
    rendering a machine completely useless.

    The patch has been reviewed for its potential impacts to other LinuxKPI
    parts and our existing DRM kmods' code.  In particular, there is no
    other user of __GFP_NORETRY/GFP_NORETRY with Linux's alloc_pages*()
    functions in our tree or DRM kmod ports.

    It has also been tested extensively, by me for months against 14-STABLE
    and sporadically on -CURRENT on a RX580, and by several others as
    reported below and as is visible in more details in the quoted bugzilla
    PR and in the initial drm-kmod issue at
    https://github.com/freebsd/drm-kmod/issues/302, on a variety of other
    AMD GPUs (several RX580, RX570, Radeon Pro WX5100, Green Sardine 5600G,
    Ryzen 9 4900H with embedded Renoir).]

    PR:             277476
    Reported by:    Josef 'Jeff' Sipek <jeffpc@josefsipek.net>
    Reviewed by:    olce
    Tested by:      many (olce, Pierre Pronchery, Evgenii Khramtsov, chaplina, rk)
    MFC after:      2 weeks
    Relnotes:       yes
    Sponsored by:   The FreeBSD Foundation (review and part of testing)

 sys/compat/linuxkpi/common/include/linux/gfp.h | 4 ++--
 sys/compat/linuxkpi/common/src/linux_page.c    | 3 ++-
 2 files changed, 4 insertions(+), 3 deletions(-)
Comment 26 commit-hook freebsd_committer freebsd_triage 2025-04-08 13:42:03 UTC
A commit in branch stable/14 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=831e6fb0baf67c2421abb50b6a14da9e71c183bb

commit 831e6fb0baf67c2421abb50b6a14da9e71c183bb
Author:     Mathieu <sigsys@gmail.com>
AuthorDate: 2024-11-14 00:24:02 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-04-08 13:38:29 +0000

    LinuxKPI: make linux_alloc_pages() honor __GFP_NORETRY

    This is to fix slowdowns with drm-kmod that get worse over time as
    physical memory become more fragmented (and probably also depending on
    other factors).

    Based on information posted in this bug report:
    https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=277476

    By default, linux_alloc_pages() retries failed allocations by calling
    vm_page_reclaim_contig() to attempt to free contiguous physical memory
    pages. vm_page_reclaim_contig() does not always succeed and calling it
    can be very slow even when it fails. When physical memory is very
    fragmented, vm_page_reclaim_contig() can end up being called (and
    failing) after every allocation attempt. This could cause very
    noticeable graphical desktop hangs (which could last seconds).

    The drm-kmod code in question attempts to allocate multiple contiguous
    pages at once but does not actually require them to be contiguous. It
    can fallback to doing multiple smaller allocations when larger
    allocations fail. It passes alloc_pages() the __GFP_NORETRY flag in this
    case.

    This patch makes linux_alloc_pages() fail early (without retrying) when
    this flag is passed.

    [olce: The problem this patch fixes is longer and longer GUI freezes as
    a machine's memory gets filled and becomes fragmented, when using amdgpu
    from DRM kmod 5.15 and DRM kmod 6.1 (DRM kmod 5.10 is unaffected; newer
    Linux kernel introduced an "optimization" by which a pool of pages is
    filled preferentially with contiguous pages, which triggered the problem
    for us).  The original commit message above evokes freezes lasting
    seconds, but I occasionally witnessed some lasting tens of minutes,
    rendering a machine completely useless.

    The patch has been reviewed for its potential impacts to other LinuxKPI
    parts and our existing DRM kmods' code.  In particular, there is no
    other user of __GFP_NORETRY/GFP_NORETRY with Linux's alloc_pages*()
    functions in our tree or DRM kmod ports.

    It has also been tested extensively, by me for months against 14-STABLE
    and sporadically on -CURRENT on a RX580, and by several others as
    reported below and as is visible in more details in the quoted bugzilla
    PR and in the initial drm-kmod issue at
    https://github.com/freebsd/drm-kmod/issues/302, on a variety of other
    AMD GPUs (several RX580, RX570, Radeon Pro WX5100, Green Sardine 5600G,
    Ryzen 9 4900H with embedded Renoir).]

    PR:             277476
    Reported by:    Josef 'Jeff' Sipek <jeffpc@josefsipek.net>
    Reviewed by:    olce
    Tested by:      many (olce, Pierre Pronchery, Evgenii Khramtsov, chaplina, rk)
    MFC after:      2 weeks
    Relnotes:       yes
    Sponsored by:   The FreeBSD Foundation (review and part of testing)

    (cherry picked from commit 718d1928f8748fe4429c011296f94f194d63c695)

 sys/compat/linuxkpi/common/include/linux/gfp.h | 4 ++--
 sys/compat/linuxkpi/common/src/linux_page.c    | 3 ++-
 2 files changed, 4 insertions(+), 3 deletions(-)
Comment 27 sigsys 2025-07-11 05:24:41 UTC
Welp, I just upgraded to 14.3-RELEASE and I got the same memory fragmentation related slowdowns. Using drm-kmod's 6.1 branch.

It's pretty much the same problem (just happening through a different route) so I'll post it here rather than opening a new PR...

The fix seems even simpler this time:

diff --git a/sys/compat/linuxkpi/common/include/linux/slab.h b/sys/compat/linuxkpi/common/include/linux/slab.h
index f3a840d9bf4b..efa5c8cb67b3 100644
--- a/sys/compat/linuxkpi/common/include/linux/slab.h
+++ b/sys/compat/linuxkpi/common/include/linux/slab.h
@@ -45,7 +45,7 @@
 
 MALLOC_DECLARE(M_KMALLOC);
 
-#define	kvzalloc(size, flags)		kmalloc(size, (flags) | __GFP_ZERO)
+#define	kvzalloc(size, flags)		kvmalloc(size, (flags) | __GFP_ZERO)
 #define	kvcalloc(n, size, flags)	kvmalloc_array(n, size, (flags) | __GFP_ZERO)
 #define	kzalloc(size, flags)		kmalloc(size, (flags) | __GFP_ZERO)
 #define	kzalloc_node(size, flags, node)	kmalloc_node(size, (flags) | __GFP_ZERO, node)

amdgpu's dc_create_state() is making large-ish (~145KB) allocations with kvzalloc() which lead to the same slowdowns in vm_page_reclaim_contig_domain_ext().

There's also a kzalloc() call in amdgpu's bw_calcs() function asking for 14400 bytes that possibly could be turned into a kvzalloc().
Comment 28 Olivier Certner freebsd_committer freebsd_triage 2025-07-11 06:25:59 UTC
Hi, yes, I'm working on this, and have this fix locally and a few other fixes that seem to considerably improve the situation, although I'm not yet sure we're finally getting to the bottom of this.  They will be posted for review soon.
Comment 29 Olivier Certner freebsd_committer freebsd_triage 2025-07-11 07:42:33 UTC
Please see D51246 to D51251 and D51253.
Comment 30 Mark Linimon freebsd_committer freebsd_triage 2025-07-20 21:22:12 UTC
^Triage: note that the following have been committed:

D51246
D51247
D51248
D51253

but not yet

D51249
D51250
D51251
Comment 31 commit-hook freebsd_committer freebsd_triage 2025-07-28 13:32:52 UTC
A commit in branch stable/14 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=8e1b8ae7298fc8ec91f6d7dc040341343c894e57

commit 8e1b8ae7298fc8ec91f6d7dc040341343c894e57
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-07 13:28:21 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-07-28 13:28:50 +0000

    LinuxKPI: Have kvzalloc() rely on kvmalloc(), not kmalloc()

    Since commit 19df0c5abcb9d4e9 ("LinuxKPI: make __kmalloc() play by the
    rules"), kmalloc() systematically allocates contiguous physical memory,
    as it should.  However, kvzalloc() was left defined in terms of
    kmalloc(), which makes it allocate contiguous physical memory too.  This
    is a too stringent restriction, as kvzalloc() is supposed to be a simple
    page-zeroing wrapper around kvmalloc().

    According to Linux's documentation ("memory-allocation.rst"), kvmalloc()
    first tries to allocate contiguous memory, falling back to
    non-contiguous one if that fails.  Thus, callers are already supposed to
    deal with the possibility of non-contiguous memory being returned.

    Reviewed by:    bz
    Fixes:          19df0c5abcb9 ("LinuxKPI: make __kmalloc() play by the rules")
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51247

    (cherry picked from commit 986edb19a49c7d7d3050c759d9b0826283492ebf)

    Forgotten on commit to main/-CURRENT:
    PR:             277476

 sys/compat/linuxkpi/common/include/linux/slab.h | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)
Comment 32 commit-hook freebsd_committer freebsd_triage 2025-07-28 13:32:56 UTC
A commit in branch stable/14 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=0e9a3bdeb7376bf04cf21830c4e33aa7efe13961

commit 0e9a3bdeb7376bf04cf21830c4e33aa7efe13961
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-08 16:17:30 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-07-28 13:28:51 +0000

    pmap: Degrade pmap_page_set_attr*() into a no-op on same attribute

    For 32-bit arm, move the no-op test that was already in place at start
    of the function so that it stays first even if the '#if 0' block around
    the call to sf_buf_invalidate_cache() is uncommented at some point (if
    ever).

    Reviewed by:    jeffpc_josefsipek.net, kib
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51253

    (cherry picked from commit 9d1f3ce79d85855399663e3977766ec46f28cadd)

    Forgotten on commit to main/-CURRENT:
    PR:             277476

 sys/amd64/amd64/pmap.c      |  5 +++++
 sys/arm/arm/pmap-v6.c       | 32 +++++++++++++++-----------------
 sys/arm64/arm64/pmap.c      |  2 ++
 sys/i386/i386/pmap.c        |  2 ++
 sys/powerpc/aim/mmu_oea.c   |  3 +++
 sys/powerpc/aim/mmu_oea64.c |  3 +++
 sys/powerpc/aim/mmu_radix.c |  4 ++++
 sys/riscv/riscv/pmap.c      |  2 ++
 8 files changed, 36 insertions(+), 17 deletions(-)
Comment 33 commit-hook freebsd_committer freebsd_triage 2025-07-28 13:32:58 UTC
A commit in branch stable/14 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=d738ed04dbff78243617ac341486b86486e9c8be

commit d738ed04dbff78243617ac341486b86486e9c8be
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-07 12:27:48 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-07-28 13:28:49 +0000

    LinuxKPI: alloc_pages(): Don't reclaim on __GFP_NORETRY

    Pass VM_ALLOC_NORECLAIM to vm_page_alloc_noobj_contig() so that it
    avoids reclaiming (currently, calling vm_reserv_reclaim_contig()).

    According to Linux's documentation, __GFP_NORETRY should not cause any
    "disruptive reclaim".  alloc_pages() is called a lot from the amdgpu DRM
    driver via ttm_pool_alloc(), which tries to allocate pages of the
    highest order first and fallback to lower order pages (as allocating
    contiguous physical pages is in fact not a requirement).  This process
    relies on failing fast, as requested by __GFP_NORETRY.  See also related
    commit 718d1928f874 ("LinuxKPI: make linux_alloc_pages() honor
    __GFP_NORETRY").

    Reviewed by:    jeffpc_josefsipek.net, bz
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51246

    (cherry picked from commit 4ca9190251bbd00c928a3cba54712c3ec25e9e26)

    Forgotten on commit to main/-CURRENT:
    PR:             277476

 sys/compat/linuxkpi/common/src/linux_page.c | 5 +++++
 1 file changed, 5 insertions(+)
Comment 34 Ivan Rozhuk 2025-08-29 23:56:15 UTC
Looks like fixed for me, thanks!
Comment 35 Olivier Certner freebsd_committer freebsd_triage 2025-09-01 15:36:40 UTC
(In reply to Ivan Rozhuk from comment #34)

What did you test? -CURRENT with only the committed patches? All patches posted to reviews?
Comment 36 Ivan Rozhuk 2025-09-01 15:43:18 UTC
(In reply to Olivier Certner from comment #35)
stable/14, patches already back ported.
Comment 37 commit-hook freebsd_committer freebsd_triage 2025-09-09 07:58:43 UTC
A commit in branch main references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=637d9858e6a8b4a8a3ee4dd80743a58bde4cbd68

commit 637d9858e6a8b4a8a3ee4dd80743a58bde4cbd68
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-08 12:28:31 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-09 07:56:51 +0000

    vm_domainset: Refactor iterators, multiple fixes

    vm_domainset_iter_first() would not check if the initial domain selected
    by the policy was effectively valid (i.e., allowed by the domainset and
    not marked as ignored by vm_domainset_iter_ignore()).  It would just try
    to skip it if it had less pages than 'free_min', and would not take into
    account the possibility of no domains being valid.

    Factor out code that logically belongs to the iterator machinery and is
    not tied to how allocations (or impossibility thereof) are to be
    handled.  This allows to remove duplicated code between
    vm_domainset_iter_page() and vm_domainset_iter_policy(), and between
    vm_domainset_iter_page_init() and _vm_domainset_iter_policy_init().
    This also allows to remove the 'pages' parameter from
    vm_domainset_iter_page_init().

    This also makes the two-phase logic clearer, revealing an inconsistency
    between setting 'di_minskip' to true in vm_domainset_iter_init()
    (implying that, in the case of waiting allocations, further attempts
    after the first sleep should just allocate for the first domain,
    regardless of their situation with respect to their 'free_min') and
    trying to skip the first domain if it has too few pages in
    vm_domainset_iter_page_init() and _vm_domainset_iter_policy_init().  Fix
    this inconsistency by resetting 'di_minskip' to 'true' in
    vm_domainset_iter_first() instead so that, after each vm_wait_doms()
    (waiting allocations that could not be satisfied immediately), we again
    start with only the domains that have more than 'free_min' pages.

    While here, fix the minor quirk that the round-robin policy would start
    with the domain after the one pointed to by the initial value of
    'di_iter' (this just affects the case of resetting '*di_iter', and would
    not cause domain skips in other circumstances, i.e., for waiting
    allocations that actually wait or at each subsequent new iterator
    creation with same iteration index storage).

    PR:             277476
    Tested by:      Kenneth Raplee <kenrap_kennethraplee.com>
    Fixes:          7b11a4832691 ("Add files for r327895")
    Fixes:          e5818a53dbd2 ("Implement several enhancements to NUMA policies.")
    Fixes:          23984ce5cd24 ("Avoid resource deadlocks when one domain has exhausted its memory."...)
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51251

 sys/kern/kern_malloc.c |  11 ++-
 sys/vm/uma_core.c      |  10 ++-
 sys/vm/vm_domainset.c  | 238 +++++++++++++++++++++++++++++--------------------
 sys/vm/vm_domainset.h  |   9 +-
 sys/vm/vm_glue.c       |   2 +-
 sys/vm/vm_kern.c       |  12 ++-
 sys/vm/vm_page.c       |  21 +++--
 7 files changed, 181 insertions(+), 122 deletions(-)
Comment 38 commit-hook freebsd_committer freebsd_triage 2025-09-09 07:58:47 UTC
A commit in branch main references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=d0b691a7c1aacf5a3f5ee6fc53f08563744d7203

commit d0b691a7c1aacf5a3f5ee6fc53f08563744d7203
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-07 20:37:14 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-09 07:56:50 +0000

    vm_domainset: Simplify vm_domainset_iter_next()

    As we are now visiting each domain only once, the test in
    vm_domainset_iter_prefer() about skipping the preferred domain (the one
    initially visited for policy DOMAINSET_POLICY_PREFER) becomes redundant.
    Removing it makes this function essentially the same as
    vm_domainset_iter_rr().

    Thus, remove vm_domainset_iter_prefer().  This makes all policies behave
    the same in vm_domainset_iter_next().

    No functional change (intended).

    PR:             277476
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51250

 sys/vm/vm_domainset.c | 32 ++------------------------------
 1 file changed, 2 insertions(+), 30 deletions(-)
Comment 39 commit-hook freebsd_committer freebsd_triage 2025-09-09 07:58:50 UTC
A commit in branch main references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=d440953942372ca275d0743a6e220631bde440ee

commit d440953942372ca275d0743a6e220631bde440ee
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-07 20:29:12 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-09 07:56:45 +0000

    vm_domainset: Only probe domains once when iterating, instead of up to 4 times

    Because of the 'di_minskip' logic, which resets the initial domain, an
    iterator starts by considering only domains that have more than
    'free_min' pages in a first phase, and then all domains in a second one.
    Non-"underpaged" domains are thus examined twice, even if the allocation
    can't succeed.

    Re-scanning the same domains twice just wastes time, as allocation
    attempts that must not wait may rely on failing sooner and those that
    must will loop anyway (a domain previously scanned twice has more pages
    than 'free_min' and consequently vm_wait_doms() will just return
    immediately).

    Additionally, the DOMAINSET_POLICY_FIRSTTOUCH policy would aggravate
    this situation by reexamining the current domain again at the end of
    each phase.  In the case of a single domain, this means doubling again
    the number of times domain 0 is probed.

    Implementation consists in adding two 'domainset_t' to 'struct
    vm_domainset_iter' (and removing the 'di_n' counter).  The first,
    'di_remain_mask', contains domains still to be explored in the current
    phase, the first phase concerning only domains with more pages than
    'free_min' ('di_minskip' true) and the second one concerning only
    domains previously under 'free_min' ('di_minskip' false).  The second,
    'di_min_mask', holds the domains with less pages than 'free_min'
    encountered during the first phase, and serves as the reset value for
    'di_remain_mask' when transitioning to the second phase.

    PR:             277476
    Fixes:          e5818a53dbd2 ("Implement several enhancements to NUMA policies.")
    Fixes:          23984ce5cd24 ("Avoid resource deadlocks when one domain has exhausted its memory."...)
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51249

 sys/vm/vm_domainset.c | 53 ++++++++++++++++++++++++++++++---------------------
 sys/vm/vm_domainset.h |  6 +++++-
 2 files changed, 36 insertions(+), 23 deletions(-)
Comment 40 Olivier Certner freebsd_committer freebsd_triage 2025-09-09 08:39:59 UTC
Will close this once the last commits have been merged to stable/15 and stable/14.
Comment 41 commit-hook freebsd_committer freebsd_triage 2025-09-19 10:08:03 UTC
A commit in branch stable/15 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=83ad6d8d8eee49223679ce9e32860dcd0ea26434

commit 83ad6d8d8eee49223679ce9e32860dcd0ea26434
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-08 12:28:31 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-19 10:07:04 +0000

    vm_domainset: Refactor iterators, multiple fixes

    vm_domainset_iter_first() would not check if the initial domain selected
    by the policy was effectively valid (i.e., allowed by the domainset and
    not marked as ignored by vm_domainset_iter_ignore()).  It would just try
    to skip it if it had less pages than 'free_min', and would not take into
    account the possibility of no domains being valid.

    Factor out code that logically belongs to the iterator machinery and is
    not tied to how allocations (or impossibility thereof) are to be
    handled.  This allows to remove duplicated code between
    vm_domainset_iter_page() and vm_domainset_iter_policy(), and between
    vm_domainset_iter_page_init() and _vm_domainset_iter_policy_init().
    This also allows to remove the 'pages' parameter from
    vm_domainset_iter_page_init().

    This also makes the two-phase logic clearer, revealing an inconsistency
    between setting 'di_minskip' to true in vm_domainset_iter_init()
    (implying that, in the case of waiting allocations, further attempts
    after the first sleep should just allocate for the first domain,
    regardless of their situation with respect to their 'free_min') and
    trying to skip the first domain if it has too few pages in
    vm_domainset_iter_page_init() and _vm_domainset_iter_policy_init().  Fix
    this inconsistency by resetting 'di_minskip' to 'true' in
    vm_domainset_iter_first() instead so that, after each vm_wait_doms()
    (waiting allocations that could not be satisfied immediately), we again
    start with only the domains that have more than 'free_min' pages.

    While here, fix the minor quirk that the round-robin policy would start
    with the domain after the one pointed to by the initial value of
    'di_iter' (this just affects the case of resetting '*di_iter', and would
    not cause domain skips in other circumstances, i.e., for waiting
    allocations that actually wait or at each subsequent new iterator
    creation with same iteration index storage).

    PR:             277476
    Tested by:      Kenneth Raplee <kenrap_kennethraplee.com>
    Fixes:          7b11a4832691 ("Add files for r327895")
    Fixes:          e5818a53dbd2 ("Implement several enhancements to NUMA policies.")
    Fixes:          23984ce5cd24 ("Avoid resource deadlocks when one domain has exhausted its memory."...)
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51251

    (cherry picked from commit 637d9858e6a8b4a8a3ee4dd80743a58bde4cbd68)

 sys/kern/kern_malloc.c |  11 ++-
 sys/vm/uma_core.c      |  10 ++-
 sys/vm/vm_domainset.c  | 238 +++++++++++++++++++++++++++++--------------------
 sys/vm/vm_domainset.h  |   9 +-
 sys/vm/vm_glue.c       |   2 +-
 sys/vm/vm_kern.c       |  12 ++-
 sys/vm/vm_page.c       |  21 +++--
 7 files changed, 181 insertions(+), 122 deletions(-)
Comment 42 commit-hook freebsd_committer freebsd_triage 2025-09-19 10:08:07 UTC
A commit in branch stable/15 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=9b8f6e2d759a357701e641816197fa2fe0386fb5

commit 9b8f6e2d759a357701e641816197fa2fe0386fb5
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-07 20:37:14 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-19 10:07:03 +0000

    vm_domainset: Simplify vm_domainset_iter_next()

    As we are now visiting each domain only once, the test in
    vm_domainset_iter_prefer() about skipping the preferred domain (the one
    initially visited for policy DOMAINSET_POLICY_PREFER) becomes redundant.
    Removing it makes this function essentially the same as
    vm_domainset_iter_rr().

    Thus, remove vm_domainset_iter_prefer().  This makes all policies behave
    the same in vm_domainset_iter_next().

    No functional change (intended).

    PR:             277476
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51250

    (cherry picked from commit d0b691a7c1aacf5a3f5ee6fc53f08563744d7203)

 sys/vm/vm_domainset.c | 32 ++------------------------------
 1 file changed, 2 insertions(+), 30 deletions(-)
Comment 43 commit-hook freebsd_committer freebsd_triage 2025-09-19 10:08:09 UTC
A commit in branch stable/15 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=da257e519bc068041f5913c1d0a8d0907520f78d

commit da257e519bc068041f5913c1d0a8d0907520f78d
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-07 20:29:12 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-19 10:06:55 +0000

    vm_domainset: Only probe domains once when iterating, instead of up to 4 times

    Because of the 'di_minskip' logic, which resets the initial domain, an
    iterator starts by considering only domains that have more than
    'free_min' pages in a first phase, and then all domains in a second one.
    Non-"underpaged" domains are thus examined twice, even if the allocation
    can't succeed.

    Re-scanning the same domains twice just wastes time, as allocation
    attempts that must not wait may rely on failing sooner and those that
    must will loop anyway (a domain previously scanned twice has more pages
    than 'free_min' and consequently vm_wait_doms() will just return
    immediately).

    Additionally, the DOMAINSET_POLICY_FIRSTTOUCH policy would aggravate
    this situation by reexamining the current domain again at the end of
    each phase.  In the case of a single domain, this means doubling again
    the number of times domain 0 is probed.

    Implementation consists in adding two 'domainset_t' to 'struct
    vm_domainset_iter' (and removing the 'di_n' counter).  The first,
    'di_remain_mask', contains domains still to be explored in the current
    phase, the first phase concerning only domains with more pages than
    'free_min' ('di_minskip' true) and the second one concerning only
    domains previously under 'free_min' ('di_minskip' false).  The second,
    'di_min_mask', holds the domains with less pages than 'free_min'
    encountered during the first phase, and serves as the reset value for
    'di_remain_mask' when transitioning to the second phase.

    PR:             277476
    Fixes:          e5818a53dbd2 ("Implement several enhancements to NUMA policies.")
    Fixes:          23984ce5cd24 ("Avoid resource deadlocks when one domain has exhausted its memory."...)
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51249

    (cherry picked from commit d440953942372ca275d0743a6e220631bde440ee)

 sys/vm/vm_domainset.c | 53 ++++++++++++++++++++++++++++++---------------------
 sys/vm/vm_domainset.h |  6 +++++-
 2 files changed, 36 insertions(+), 23 deletions(-)
Comment 44 commit-hook freebsd_committer freebsd_triage 2025-09-19 11:43:25 UTC
A commit in branch stable/14 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=2c2bb7913d478889da548fe9837dc9b32727a81d

commit 2c2bb7913d478889da548fe9837dc9b32727a81d
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-08 12:28:31 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-19 11:41:49 +0000

    vm_domainset: Refactor iterators, multiple fixes

    vm_domainset_iter_first() would not check if the initial domain selected
    by the policy was effectively valid (i.e., allowed by the domainset and
    not marked as ignored by vm_domainset_iter_ignore()).  It would just try
    to skip it if it had less pages than 'free_min', and would not take into
    account the possibility of no domains being valid.

    Factor out code that logically belongs to the iterator machinery and is
    not tied to how allocations (or impossibility thereof) are to be
    handled.  This allows to remove duplicated code between
    vm_domainset_iter_page() and vm_domainset_iter_policy(), and between
    vm_domainset_iter_page_init() and _vm_domainset_iter_policy_init().
    This also allows to remove the 'pages' parameter from
    vm_domainset_iter_page_init().

    This also makes the two-phase logic clearer, revealing an inconsistency
    between setting 'di_minskip' to true in vm_domainset_iter_init()
    (implying that, in the case of waiting allocations, further attempts
    after the first sleep should just allocate for the first domain,
    regardless of their situation with respect to their 'free_min') and
    trying to skip the first domain if it has too few pages in
    vm_domainset_iter_page_init() and _vm_domainset_iter_policy_init().  Fix
    this inconsistency by resetting 'di_minskip' to 'true' in
    vm_domainset_iter_first() instead so that, after each vm_wait_doms()
    (waiting allocations that could not be satisfied immediately), we again
    start with only the domains that have more than 'free_min' pages.

    While here, fix the minor quirk that the round-robin policy would start
    with the domain after the one pointed to by the initial value of
    'di_iter' (this just affects the case of resetting '*di_iter', and would
    not cause domain skips in other circumstances, i.e., for waiting
    allocations that actually wait or at each subsequent new iterator
    creation with same iteration index storage).

    PR:             277476
    Tested by:      Kenneth Raplee <kenrap_kennethraplee.com>
    Fixes:          7b11a4832691 ("Add files for r327895")
    Fixes:          e5818a53dbd2 ("Implement several enhancements to NUMA policies.")
    Fixes:          23984ce5cd24 ("Avoid resource deadlocks when one domain has exhausted its memory."...)
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51251

    (cherry picked from commit 637d9858e6a8b4a8a3ee4dd80743a58bde4cbd68)

 sys/kern/kern_malloc.c |  11 ++-
 sys/vm/uma_core.c      |  10 ++-
 sys/vm/vm_domainset.c  | 234 +++++++++++++++++++++++++++++--------------------
 sys/vm/vm_domainset.h  |   6 +-
 sys/vm/vm_kern.c       |  12 ++-
 sys/vm/vm_page.c       |  20 +++--
 6 files changed, 177 insertions(+), 116 deletions(-)
Comment 45 commit-hook freebsd_committer freebsd_triage 2025-09-19 11:43:30 UTC
A commit in branch stable/14 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=b5834753d330f201195602d039633be8d811e204

commit b5834753d330f201195602d039633be8d811e204
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-07 20:29:12 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-19 11:41:41 +0000

    vm_domainset: Only probe domains once when iterating, instead of up to 4 times

    Because of the 'di_minskip' logic, which resets the initial domain, an
    iterator starts by considering only domains that have more than
    'free_min' pages in a first phase, and then all domains in a second one.
    Non-"underpaged" domains are thus examined twice, even if the allocation
    can't succeed.

    Re-scanning the same domains twice just wastes time, as allocation
    attempts that must not wait may rely on failing sooner and those that
    must will loop anyway (a domain previously scanned twice has more pages
    than 'free_min' and consequently vm_wait_doms() will just return
    immediately).

    Additionally, the DOMAINSET_POLICY_FIRSTTOUCH policy would aggravate
    this situation by reexamining the current domain again at the end of
    each phase.  In the case of a single domain, this means doubling again
    the number of times domain 0 is probed.

    Implementation consists in adding two 'domainset_t' to 'struct
    vm_domainset_iter' (and removing the 'di_n' counter).  The first,
    'di_remain_mask', contains domains still to be explored in the current
    phase, the first phase concerning only domains with more pages than
    'free_min' ('di_minskip' true) and the second one concerning only
    domains previously under 'free_min' ('di_minskip' false).  The second,
    'di_min_mask', holds the domains with less pages than 'free_min'
    encountered during the first phase, and serves as the reset value for
    'di_remain_mask' when transitioning to the second phase.

    PR:             277476
    Fixes:          e5818a53dbd2 ("Implement several enhancements to NUMA policies.")
    Fixes:          23984ce5cd24 ("Avoid resource deadlocks when one domain has exhausted its memory."...)
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51249

    (cherry picked from commit d440953942372ca275d0743a6e220631bde440ee)

 sys/vm/vm_domainset.c | 53 ++++++++++++++++++++++++++++++---------------------
 sys/vm/vm_domainset.h |  6 +++++-
 2 files changed, 36 insertions(+), 23 deletions(-)
Comment 46 commit-hook freebsd_committer freebsd_triage 2025-09-19 11:43:33 UTC
A commit in branch stable/14 references this bug:

URL: https://cgit.FreeBSD.org/src/commit/?id=0206ec4350ca385a8c29f410dc6814b5517dafc5

commit 0206ec4350ca385a8c29f410dc6814b5517dafc5
Author:     Olivier Certner <olce@FreeBSD.org>
AuthorDate: 2025-07-07 20:37:14 +0000
Commit:     Olivier Certner <olce@FreeBSD.org>
CommitDate: 2025-09-19 11:41:42 +0000

    vm_domainset: Simplify vm_domainset_iter_next()

    As we are now visiting each domain only once, the test in
    vm_domainset_iter_prefer() about skipping the preferred domain (the one
    initially visited for policy DOMAINSET_POLICY_PREFER) becomes redundant.
    Removing it makes this function essentially the same as
    vm_domainset_iter_rr().

    Thus, remove vm_domainset_iter_prefer().  This makes all policies behave
    the same in vm_domainset_iter_next().

    No functional change (intended).

    PR:             277476
    MFC after:      10 days
    Sponsored by:   The FreeBSD Foundation
    Differential Revision:  https://reviews.freebsd.org/D51250

    (cherry picked from commit d0b691a7c1aacf5a3f5ee6fc53f08563744d7203)

 sys/vm/vm_domainset.c | 32 ++------------------------------
 1 file changed, 2 insertions(+), 30 deletions(-)
Comment 47 Olivier Certner freebsd_committer freebsd_triage 2025-09-19 22:01:36 UTC
With all the fixes above, I'm still seeing slowdowns once all the machine's memory is fragmented, but I had to wait for a much longer time for the situation to degrade.

Another fix is needed:
https://github.com/freebsd/drm-kmod/pull/377
(It's basically what sigsys@ mentioned above, except I don't find it optional. :-))

With this I've yet to experience slowdowns with DRM 6.1.  Observing free memory and its fragmentation shows the latter has diminished tremendously after this last fix even after a while, giving confidence that nothing is trying to allocate non-trivial amount of memory (> PAGE_SIZE) in physically contiguous chunks in the fast path.

So we are perhaps at the bottom of this now.  FreeBSD apparently does not defragment enough to sustain frequent contiguous allocations (probably Linux does, assuming nobody complains about slowdowns there?), but as a matter of fact these are not really needed.  In new DRM code, we'll have to watch out for new allocations of such kind, which would degrade performance again until patched.
Comment 48 Olivier Certner freebsd_committer freebsd_triage 2025-09-24 11:21:31 UTC
(In reply to Olivier Certner from comment #47)

The above-mentioned pull request was merged, as well as companion ones (backports to 5.10, 5.15, 6.1 and 6.6).

So really, you should not observe any more slowdowns (at least due to physically-contiguous allocations) on any versions built from recent stable/14 and stable/15, whichever DRM kmod version you are using.  Please make sure to recompile both your kernels and the your drm-kmod module.

Release-wise, the fixes will be in 15.0, and 14.4 (first semester 2026).  We may want to issue ENs for 14 though.

Due to the nature of the problem and fixes, we might still experience regressions as drm-kmod but also our LinuxKPI are updated (we had such a regression for 14.3, annihilating the benefit of my initial fixes over 14.2 at the beginning of the year), re-introducing physically-contiguous allocations that are not really needed in hot paths.

So please test and report as things evolve.

Thanks and regards.
Comment 49 Olivier Certner freebsd_committer freebsd_triage 2025-10-02 08:44:11 UTC
FYI: As for the DRM changes, they have landed in the drm-kmod GitHub repo, but not yet in releases, so are not currently in any of our DRM ports.  I'm working with DRM guys to ensure that these will be updated before the 15.0 release, despite 2025Q4 having been branched yesterday.