MemTest86 Pro 11.5/11.6 causes UEFI Fault (Dell PowerEdge R7725 - AMD EPYC 9335)

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • hdwr@edgeam.am
    Junior Member
    • Sep 2026
    • 7

    #1

    MemTest86 Pro 11.5/11.6 causes UEFI Fault (Dell PowerEdge R7725 - AMD EPYC 9335)

    Здравствуйте,

    мы столкнулись с периодически возникающей, но воспроизводимой проблемой при запуске MemTest86 Pro на серверах Dell PowerEdge R7725, оснащенных процессорами AMD EPYC 9335.
    Мы воспроизвели проблему на двух разных платформах Dell PowerEdge R7725 SFF8, обе с процессорами AMD EPYC 9335 (32-ядерные).
    Мы протестировали MemTest86 Pro версий 11.5 и 11.6, и проблема возникает в обеих версиях.
    Все микропрограммы сервера обновлены до последних версий, доступных на данный момент у Dell, включая BIOS, iDRAC/Lifecycle Controller и другие системные микропрограммы.

    Основная конфигурация R7725:
    Сервер: Dell PowerEdge R7725 SFF8
    Процессор: 2 × AMD EPYC 9335, 3,0 ГГц, по 32 ядра каждый
    Память: 16 × 32 ГБ DDR5 RDIMM 6400 МТ/с Dell (всего 512 ГБ)
    Видеокарта: 2 × NVIDIA RTX PRO 6000 BSE 96 ГБ
    Сеть: Broadcom 57414 двухпортовый 25GbE SFP28 OCP
    Хранение данных: 2 × 1,92 ТБ NVMe E3.s 2,5" RI Dell SSD
    Блок питания: Dell Dual Hot-Plug 3200 Вт
    Охлаждение: радиаторы для процессоров Dell R7725 и высокопроизводительные вентиляторы 2U Platinum
    Диагностика оборудования уже проведена.

    Мы провели обширное тестирование изоляции оборудования и замены компонентов.
    Оба процессора AMD EPYC 9335 Процессоры были заменены на другие/новые ЦП, но сбой MemTest86/UEFI по-прежнему происходит.
    Мы также протестировали сервер с различными модулями памяти, и проблема осталась.

    Чтобы исключить графические процессоры и устройства PCIe как возможную причину, мы также протестировали сервер с полностью удаленными графическими процессорами NVIDIA RTX PRO 6000. Тот же сбой UEFI по-прежнему происходит.
    Кроме того, у нас есть вторая/запасная платформа Dell PowerEdge R7725, и мы также провели те же тесты изоляции и замены компонентов на этой системе. Тот же сбой MemTest86/UEFI воспроизводится на второй платформе R7725.
    Таким образом, проблема не ограничивается какой-либо конкретной материнской платой, комплектом ЦП, комплектом памяти или конфигурацией графического процессора.

    MemTest86 не сообщает о каких-либо ошибках памяти перед сбоем. Вместо этого вся среда UEFI в конечном итоге дает сбой, и сервер останавливается со следующей ошибкой:
    PowerEdge R7725 – BIOS
    Требуется перезагрузка системы. Система обнаружила исключение во время UEFI. Предзагрузочная среда.
    Тип: Общая ошибка защиты (13)
    Источник: Программное обеспечение (UEFI0011) на BSP
    Трассировка стека на экране исключения включает:
    DxeCore.efi
    HiiDatabase.efi
    DellEarlyVideoSplashDxe.efi

    В журнале жизненного цикла iDRAC записано только:
    BIOS системы остановлен.

    К сожалению, нам не удалось получить журналы MemTest86 после сбоя. Когда возникает проблема, сама среда BIOS/UEFI останавливается, и результаты/журналы тестирования не сохраняются.

    Сбой носит периодический характер и не происходит в одной и той же точке при каждом запуске. Иногда MemTest86 завершается успешно без ошибок. Однако при последующих запусках или дополнительных проходах система может в конечном итоге аварийно завершиться с ошибкой общей защиты, показанной выше.

    Во всех случаях, когда нам удалось воспроизвести проблему, сбой происходил примерно через 17–29 часов непрерывной работы MemTest86.
    До сбоя MemTest86 не сообщает об ошибках памяти.

    В качестве дополнительной проверки мы провели расширенное стресс-тестирование под Linux с использованием stress-ng.
    Системы R7725 остаются стабильными при стресс-тестировании под Linux, и мы не наблюдали там аналогичного сбоя.

    Для сравнения мы также протестировали MemTest86 Pro 11.6 на Dell PowerEdge R7625 предыдущего поколения с AMD EPYC 9224 (Genoa).
    Мы не наблюдали этой проблемы на R7625 — MemTest86 успешно работает без сбоев UEFI.
    На данный момент проблема воспроизведена только на Dell PowerEdge R7725 с AMD EPYC 9335 (Turin) на двух разных платформах R7725, в то время как на предыдущем поколении R7625 с AMD EPYC 9224 (Genoa) аналогичного поведения не наблюдается.

    Следует учитывать, что:
    проблема воспроизведена на двух разных платформах Dell PowerEdge R7725;
    затронуты как MemTest86 Pro 11.5, так и 11.6;
    процессоры были заменены;
    были протестированы различные модули памяти;
    системы тестировались без установленных графических процессоров NVIDIA;
    аналогичное тестирование с заменой/изоляцией было проведено на второй/запасной платформе R7725;
    MemTest86 иногда завершается успешно без ошибок;
    когда происходит сбой, это ошибка UEFI General Protection Fault, а не ошибка памяти MemTest86;
    сбой обычно происходит только после 17–29 часов непрерывного тестирования;
    стресс-тестирование Linux stress-ng стабильно;
    На процессорах предыдущего поколения R7625 / EPYC 9224 эта проблема отсутствует.

    Мы хотели бы определить, может ли это быть проблемой совместимости между MemTest86 и более новой платформой AMD EPYC 9005 / Turin, или же это проблема реализации UEFI на Dell PowerEdge R7725.

    Пожалуйста, сообщите:
    Известны ли какие-либо проблемы с MemTest86 11.5/11.6 на платформах AMD EPYC 9005 (Turin)?
    Тестировался ли MemTest86 специально на Dell PowerEdge R7725?
    Может ли сбой быть связан с топологией ЦП, доступом к SPD/SMBus, обработкой ECC, опросом температуры памяти или другими аппаратными запросами, выполняемыми MemTest86?
    Рекомендуете ли вы тестирование с более новой или внутренней/отладочной сборкой MemTest86, если таковая доступна?

    Прилагаю скриншоты полного сообщения об ошибке UEFI и соответствующей записи в журнале жизненного цикла iDRAC. Также могу предоставить подробную информацию о версиях прошивки, конфигурации памяти, скриншоты MemTest86 и любую дополнительную информацию о системе, необходимую для устранения неполадок.
  • David (PassMark)
    Administrator
    • Jan 2003
    • 11109

    #2
    We don't speak Russian. So just guessing from a bad translation.

    You are using an old version of Memtest86. See this page for the latest release.
    https://www.memtest86.com/whats-new.html
    But using V11.7 probably won't fix this.

    The crash is a Dell UEFI firmware exception, not a MemTest86 reported memory test failure.
    Strangely the timestamp showed this happened 5 months ago??

    Faulting instruction was in the EFI software, DxeCore.efi at address 0x02E085
    DXE = Driver Execution Environment.

    Call path was:
    DellEarlyVideoSplashDxe.efi called
    HiiDatabase.efi which called
    DxeCore.efi which crashed

    There are multiple similar crash on the Dell web site:

    https://www.dell.com/community/en/co...ccf8a8de29c346
    https://www.dell.com/community/en/co...356218a14fb217
    https://www.dell.com/community/en/co...ccf8a8de49fe86

    You might need to get Dell involved to debug their BIOS.

    Comment

    • hdwr@edgeam.am
      Junior Member
      • Sep 2026
      • 7

      #3
      Sorry, here is the original text:
      Hello,

      We are experiencing an intermittent but reproducible issue when running MemTest86 Pro on Dell PowerEdge R7725 servers equipped with AMD EPYC 9335 processors.
      We have reproduced the issue on two different Dell PowerEdge R7725 SFF8 platforms, both using AMD EPYC 9335 (32-core) processors.
      We tested both MemTest86 Pro 11.5 and 11.6, and the issue occurs with both versions.
      All server firmware is up to date with the latest versions currently available from Dell, including BIOS, iDRAC/Lifecycle Controller, and other system firmware.

      The primary R7725 configuration is:
      Server: Dell PowerEdge R7725 SFF8
      CPU: 2 × AMD EPYC 9335, 3.0 GHz, 32 cores each
      Memory: 16 × 32 GB DDR5 RDIMM 6400 MT/s Dell (512 GB total)
      GPU: 2 × NVIDIA RTX PRO 6000 BSE 96 GB
      Network: Broadcom 57414 dual-port 25GbE SFP28 OCP
      Storage: 2 × 1.92 TB NVMe E3.s 2.5" RI Dell SSD
      PSU: Dell Dual Hot-Plug 3200 W Power Supply
      Cooling: Dell R7725 CPU heatsinks and 2U High Performance Platinum fans
      Hardware troubleshooting already performed

      We have performed extensive hardware isolation and component replacement testing.
      Both AMD EPYC 9335 processors have been replaced with different/new CPUs, but the MemTest86/UEFI crash still occurs.
      We have also tested the server with different memory modules, and the issue remains.

      To exclude the GPUs and PCIe devices as a possible cause, we also tested the server with the NVIDIA RTX PRO 6000 GPUs completely removed. The same UEFI crash still occurs.
      In addition, we have a second/spare Dell PowerEdge R7725 platform, and we performed the same component isolation and replacement tests on that system as well. The same MemTest86/UEFI failure can be reproduced on the second R7725 platform.
      Therefore, the issue is not limited to one particular motherboard, CPU set, memory set, or GPU configuration.

      MemTest86 does not report any memory errors before the failure. Instead, the entire UEFI environment eventually crashes and the server halts with the following error:
      PowerEdge R7725 – BIOS
      A system restart is required. The system detected an exception during the UEFI pre-boot environment.
      Type: General Protection Fault (13)
      Source: Software (UEFI0011) on BSP
      The stack trace on the exception screen includes:
      DxeCore.efi
      HiiDatabase.efi
      DellEarlyVideoSplashDxe.efi

      The iDRAC Lifecycle Log only records:
      System BIOS has halted.

      Unfortunately, we have not been able to obtain MemTest86 logs after the crash. When the issue occurs, the BIOS/UEFI environment itself halts and the test results/logs are not saved.

      The failure is intermittent and does not occur at exactly the same point during every run. Sometimes MemTest86 completes successfully with zero errors. However, during subsequent runs or additional passes, the system may eventually crash with the General Protection Fault shown above.

      In all cases where we have reproduced the issue so far, the crash has occurred after approximately 17 to 29 hours of continuous MemTest86 operation.
      No memory errors are reported by MemTest86 before the crash.

      As an additional check, we performed extended stress testing under Linux using stress-ng.
      The R7725 systems remain stable under Linux stress testing, and we have not observed the same failure there.

      For comparison, we also tested MemTest86 Pro 11.6 on a previous-generation Dell PowerEdge R7625 with AMD EPYC 9224 (Genoa).
      We have not observed this issue on the R7625 - MemTest86 runs successfully without UEFI crashes.
      So far, the issue has been reproduced only on the Dell PowerEdge R7725 with AMD EPYC 9335 (Turin), on two separate R7725 platforms, while the previous-generation R7625 with AMD EPYC 9224 (Genoa) does not exhibit the same behavior.

      Considering that:
      the problem has been reproduced on two separate Dell PowerEdge R7725 platforms;
      both MemTest86 Pro 11.5 and 11.6 are affected;
      the processors have been replaced;
      different memory modules have been tested;
      the systems have been tested without the NVIDIA GPUs installed;
      the same isolation/replacement testing has been performed on the second/spare R7725 platform;
      MemTest86 sometimes completes successfully with zero errors;
      when the failure occurs, it is a UEFI General Protection Fault rather than a MemTest86 memory error;
      the failure typically occurs only after 17–29 hours of continuous testing;
      Linux stress-ng testing is stable;
      and a previous-generation R7625 / EPYC 9224 does not exhibit the issue,

      We would like to determine whether this could be a compatibility issue between MemTest86 and the newer AMD EPYC 9005 / Turin platform, or with the Dell PowerEdge R7725 UEFI implementation.

      Could you please advise:
      Are there any known issues with MemTest86 11.5/11.6 on AMD EPYC 9005 (Turin) platforms?
      Has MemTest86 been tested specifically on the Dell PowerEdge R7725?
      Could the crash be related to CPU topology, SPD/SMBus access, ECC handling, memory temperature polling, or other hardware interrogation performed by MemTest86?
      Would you recommend testing with a newer or internal/debug build of MemTest86, if one is available?

      I have attached screenshots of the complete UEFI exception and the corresponding iDRAC Lifecycle Log entry. I can also provide detailed firmware versions, memory configuration, MemTest86 screenshots, and any additional system information required for troubleshooting.

      Thank you.


      Thank you for the analysis.

      Regarding the timestamp: yes, the screenshot is from April 17, 2026. We have been troubleshooting this issue with the hardware supplier for several months, including replacing CPUs, memory and testing another R7725 platform.

      The issue is still reproducible. We will now repeat the test using MemTest86 Pro 11.7 on both R7725 systems.

      If the same UEFI exception occurs with 11.7, we will provide the new crash screen and timestamp as well.

      Thank you for identifying the faulting module and call path. This information will be very useful when escalating the issue to Dell.

      Comment

      • David (PassMark)
        Administrator
        • Jan 2003
        • 11109

        #4
        We have been troubleshooting this issue with the hardware supplier for several month
        Dell should have access to the source code for the their own BIOS.
        There should also be compiler map files. Using that they should be able to use to identify the exact functions that was called in DxeCore.efi, plus the call chain leading up to the crash.
        i.e. they should be able to get the exact line of code that crashed, if they know what they are doing. But this kind of expertise (with source code access) is probably hard to find in Dell. But it is a 10min job for the right person.

        With that information the root cause will probably be more obvious. e.g. out of memory, null pointer access, threading issues, etc...

        Comment

        • hdwr@edgeam.am
          Junior Member
          • Sep 2026
          • 7

          #5
          Hello,
          I have an additional finding that may be relevant to this issue.
          As mentioned previously, we started testing a previous-generation Dell PowerEdge R7625 with AMD EPYC 9224 (Genoa) for comparison with the R7725 / EPYC 9335 (Turin) systems.
          Initially, the R7625 appeared to be stable. However, after extending the test duration, we found that MemTest86 also hangs on the R7625, although the behavior is different from the R7725 UEFI General Protection Fault.
          We tested both:
          • MemTest86 Pro 11.6 — 2 attempts
          • MemTest86 Pro 11.7 — 2 attempts
          All four attempts eventually hung, at approximately the same stage of the test. Dell PowerEdge R7625 configuration
          • Dell PowerEdge R7625
          • Dell PowerEdge R7625 Motherboard V4
          • 2 × AMD EPYC 9224 2.50 GHz / 24-core
          • 24 × Samsung 32 GB DDR5 RDIMM 4800 MT/s (768 GB total)
          • 6 × Dell R7625 Very High Performance Fans
          • Standard heatsinks for 2-CPU configuration
          • Dell E810-XXV 2×25Gb SFP28 OCP 3.0 Intel
          • Intel E810-XXV 2×25Gb SFP28 LP Dell
          • Dell 5720 2×1Gb RJ-45 LOM
          • Dell Dual Hot-Plug 1400 W Power Supply
          • Dell BOSS-N1 controller with 2 × 960 GB M.2 in RAID1
          • 960 GB SATA 2.5" Dell SSD
          • Dell PERC H755
          The server firmware is also up to date. MemTest86 11.6
          In one of the reproduced hangs, the screen remained at:
          Time: 20:00:23
          The MemTest86 interface was completely stuck at this point and no further progress was observed. MemTest86 11.7
          We then repeated the test using the current MemTest86 Pro 11.7 release.
          It hung in almost exactly the same place:
          Time: 20:02:44
          Again, passes #1 and #2 completed successfully with zero memory and ECC errors before the hang.
          We have reproduced this behavior twice with 11.6 and twice with 11.7. The hangs occur at approximately the same stage, after around 20 hours of testing.
          Unlike the R7725 / EPYC 9335 issue reported earlier, the R7625 does not display the Dell UEFI General Protection Fault screen. MemTest86 simply stops making progress and remains stuck.
          So at this point we have observed two different behaviors:
          Dell R7725 / AMD EPYC 9335 (Turin):
          • MemTest86 Pro 11.5 and 11.6
          • intermittent Dell UEFI General Protection Fault
          • typically after approximately 17–29 hours
          • no MemTest86 memory errors before the crash
          • reproduced on two different R7725 platforms
          • CPUs and memory have been replaced/tested with different components
          • GPUs have been removed for testing
          • Linux stress-ng testing is stable
          Dell R7625 / AMD EPYC 9224 (Genoa):
          • MemTest86 Pro 11.6 and 11.7
          • 4 reproduced hangs in total (2 with each version)
          • approximately 20 hours into the test
          • Pass 3/4, around 85% pass progress
          • Test 12 (Random number sequence, 128-bit)
          • first two passes complete successfully
          • zero memory errors and zero ECC errors before the hang
          • no UEFI General Protection Fault; MemTest86 simply stops progressing
          I have attached screenshots from both the 11.6 and 11.7 R7625 runs.
          Given that we are now seeing reproducible long-duration issues on both AMD EPYC 9004 (Genoa) and AMD EPYC 9005 (Turin) Dell PowerEdge platforms, could there be a known limitation or compatibility issue with MemTest86 on these newer AMD EPYC generations, particularly in dual-socket systems with large DDR5 memory configurations?
          The fact that the R7625 hangs at nearly the same point in Test 12 on both MemTest86 11.6 and 11.7, after two complete error-free passes, seems particularly interesting.
          Is there any known issue involving:
          • Test 12 / Random Number Sequence (128-bit);
          • multicore execution on dual-socket AMD EPYC systems;
          • large DDR5 memory configurations;
          • NUMA/memory topology on Genoa or Turin;
          • or long-duration testing on these platforms?
          Are there any configuration options you would recommend changing (for example, CPU selection/multicore mode, Test 12 settings, SPD access, or other options) to help isolate the cause?
          If you have a debug build or additional logging method that could help determine where MemTest86 is hanging, we can test it on the R7625 and R7725 systems.
          Thank you.
          Click image for larger version

Name:	memtest-fail-11-6.png
Views:	0
Size:	104.3 KB
ID:	60741
          Click image for larger version

Name:	memtest-fail-11-7.png
Views:	0
Size:	106.1 KB
ID:	60742

          Comment

          • hdwr@edgeam.am
            Junior Member
            • Sep 2026
            • 7

            #6
            Additional test result with a smaller memory configuration on the same R7625 platform
            I have another result that may help narrow down the R7625 issue.
            We tested the same Dell PowerEdge R7625 / AMD EPYC 9224 (Genoa) platform using a smaller memory configuration with 16 GB DDR5-4800 RDIMMs (Hynix).
            With this configuration, MemTest86 Pro 11.6 completed all 4 passes successfully and did not hang.
            This suggests that the R7625 hang may be related not simply to the Genoa platform itself, but possibly to the memory capacity, number of DIMMs, DIMM topology, or the specific Samsung 32 GB RDIMM configuration.
            The fact that the smaller Hynix configuration completes MemTest86 11.6 successfully while the 24 × 32 GB Samsung configuration reproducibly hangs in approximately the same place seems significant.
            I have attached the screenshot showing the successful MemTest86 Pro 11.6 run with the smaller memory configuration.Click image for larger version

Name:	memtest-11-6-256GB.png
Views:	0
Size:	109.1 KB
ID:	60744

            Comment

            • David (PassMark)
              Administrator
              • Jan 2003
              • 11109

              #7
              To me this looks like a software bug in the UEFI BIOS.

              There is a class of bugs related to resource exhaustion. Resource exhaustion software bugs occur when a program fails to release system resources after using them, eventually consuming all available capacity and causing a system crash or freeze. When these bugs cause a crash after a predictable or specific amount of time, they are usually driven by a steady, continuous accumulation of unreleased resources during uptime. Example of these type of bugs are:
              Memory Leaks: The most common cause. A program allocates memory (RAM) but fails to free it back to the system when it is no longer needed. Over time, the system runs completely out of memory (OOM), forcing the operating system to terminate the process.
              Integer and Counter Overflows: Software often tracks time, ticks, or sequence numbers using a fixed-size integer. If the software runs continuously, the counter eventually hits its maximum value and overflows (often wrapping around to zero or a negative number), causing logic failures or immediate crashes.
              File Descriptor & Handle Leaks: Every time a program opens a file, network socket, or database connection, it uses a "handle" or "file descriptor." If the program forgets to close them, it will eventually hit the operating system's strict maximum limit, preventing any new operations and crashing.
              Thread & Process Exhaustion: Programs that spin up new background threads for tasks without properly terminating old ones will eventually exhaust the operating system's thread limit or available CPU scheduling resources.

              So having the crash happen at nearly exactly 20h each time points to a bug of these type. It also explains why all shorter tests don't crash.

              If it isn't resource exhaustion, then it might be some timer in the BIOS that expires at 20h (e.g. a watchdog timer to reset remote access). So in that case you might find BIOS always crashes at 20h even if you aren't running any memory tests. e.g. you are just sitting in the config menu.

              Comment

              Working...