I bought an Optiplex 5040, with an i5-6500TE, and 8 GB DDR3L RAM.

When I bought it, I installed Fedora Server on it. It got stuck every few days but I could never see the error. The services just stopped working, I couldn’t ssh into it, and connecting it to a monitor showed a black screen.

So, I thought let’s install Ubuntu Server, maybe Fedora isn’t compatible with all of its hardware. The same thing is happening, now, but I can see this error. Even when there’s nothing installed on it, no containers, nothing other than base packages, this happens.

I have updated the bios. I have tried setting nouveau.modeset=0 in the grub config file. I have tried disabling and enabling c-states. No luck till now.

Would really appreciate if anyone helps me with this.

UPDATE:

  • I cleaned everything and reapplied the thermal paste. I did not see any change in the thermals. It never goes over 55°C even under full load.
  • I reset the motherboard by removing that jumper thing.
  • I ran memtest86, which took over 2½ hours. It did not show any errors.
  • I ran a CPU stress test for over 15 hours, and nothing crashed.
  • I also ran the Dell’s diagnostic tool, available in the boot menu of the motherboard. The whole test took over 2 hours but did not show any errors. It tested the memory, CPU, fans, storage drives, etc.
53 points
*

Seen this before, and almost always has to do with hardware failure or bad hardware config.

Reset the BIOS/CMOS jumper on the board, go back into BIOS setup and set the proper time. Do not touch the CPU or Memory timings. Boot with the defaults and see if it still happens. Check and update the BIOS if there is a newer version as well.

Next longer steps: test memory, then stress test the CPU. I’d be shocked if it was a storage issue as I haven’t seen that be the culprit, but might was well run the long SMART tests.

permalink
report
reply
20 points

Make sure the microcode package is up to date as well

permalink
report
parent
reply
8 points

Yeah, always check all of this stuff. Server hardware gets a lot more updates than like gamer board BIOS, companies invest high millions, even low billions in this stuff and they expect problems to be address promptly for that kind of cash.

Check for any peripherals or cards, too. RAID, backplanes, networking cards; drivers, firmware, anything.

permalink
report
parent
reply
5 points

Poorly supported hardware will also do this

permalink
report
parent
reply
8 points

An Optiplex 5040 should be well and thoroughly supported for 6+y now

permalink
report
parent
reply
1 point

I did reset it. It did not help. I ran memtest86 for over 2 hours and did a CPU stress test for over 15 hours. Nothing crashed during the testing.

permalink
report
parent
reply
24 points

Had the same issues, it was heat.

Cool down your server, add a fan or a cooler…

I added a usb-powered fan sucking cooler air from outside the server area directly blowing it on the chassis.

That fixed for me.

permalink
report
reply
1 point

I cleaned everything and reapplied the thermal paste. That did not solve the problem. Also, the CPU is only of 35 watts and never goes over 55°C.

permalink
report
parent
reply
15 points

That server sounds a bit older in the teeth… Has new thermal paste been applied to the cpu? Even if the reported temps are under 90c, you might be getting hot spots causing glitches inside the package.

Worth trying a couple of different generations of kernel as well, both newer and older. You might be hitting a regression somewhere.

permalink
report
reply
2 points

I cleaned everything and reapplied the thermal paste. That did not solve the problem. Also, the CPU is only of 35 watts and never goes over 55°C.

permalink
report
parent
reply
13 points

Try running memtest86 for a few days to test memory. That is fairly easy to do though it involves booting from a flash drive. Web search should find info.

permalink
report
reply
1 point

I just ran it. It took over 2 hours to finish. Showed no errors. Is there a benefit of running it for a few days?

permalink
report
parent
reply
2 points

If the problem is intermittent then longer run has better chance of catching it, but 2 hours with no errors is a good sign, with regard to the memory.

permalink
report
parent
reply
12 points

Time to add a cron job to auto reboot it once a day

permalink
report
reply

Selfhosted

!selfhosted@lemmy.world

Create post

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don’t control.

Rules:

  1. Be civil: we’re here to support and learn from one another. Insults won’t be tolerated. Flame wars are frowned upon.

  2. No spam posting.

  3. Posts have to be centered around self-hosting. There are other communities for discussing hardware or home computing. If it’s not obvious why your post topic revolves around selfhosting, please include details to make it clear.

  4. Don’t duplicate the full text of your blog or github here. Just post the link for folks to click.

  5. Submission headline should match the article title (don’t cherry-pick information from the title to fit your agenda).

  6. No trolling.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

Community stats

  • 3.4K

    Monthly active users

  • 3.4K

    Posts

  • 77K

    Comments