Post 1 Posted by OverDrive 2008-12-24 04:23:22 UTC
GU has been unstable as hell since I installed the 2nd CPU and until I can get that fixed (only has issues under load) I'll have to go easy on game servers.
Personally, I think the box hates me - either that or loves toying with me.
Anyway - It's a cross I have to bear.
OD
Post 2 Posted by grandinferno 2008-12-24 04:29:18 UTC
Good luck with that. Somebody seems to have cursed us! :(
Post 3 Posted by Bad Crispy 2008-12-24 04:42:40 UTC
Quote from grandinferno;359537: Good luck with that. Somebody seems to have cursed us! :(
Ill blame unseen :D
However thanks OD for getting it back up ;)
Post 4 Posted by Milenko 2008-12-24 05:44:01 UTC
Yeah cheers for all your hardwork Dan ;)
Post 5 Posted by ResLo 2008-12-24 06:08:57 UTC
Phew Christmas can go ahead now....
edit - On the plus side I didn't have to suffer rec and his 10 new topics a day
Post 6 Posted by Unseen 2008-12-24 07:07:20 UTC
Quote from Bad Crispy;359539: Ill blame unseen :D
However thanks OD for getting it back up ;)
can only pick on the guy that rips your head time and time again ?? ;p
Post 7 Posted by Korny 2008-12-24 14:20:11 UTC
did the lights out card work good to help solve this problem ?
Post 8 Posted by OverDrive 2008-12-29 08:35:31 UTC
Re: The server crashes - you hardware guru's, does this look like a ram issue?
Got this from the MCE log.
MCE 0 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC dbb88c6eb92d ADDR 413916c0 Northbridge Chipkill ECC error Chipkill ECC syndrome = 2254 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node response, request didn't time out generic read mem transaction memory access, level generic' STATUS d42a400022080a13 MCGSTATUS 0 MCE 1 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC 1b1e956b6bc4d ADDR 413918a0 Northbridge Chipkill ECC error Chipkill ECC syndrome = b741 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node response, request didn't time out generic read mem transaction memory access, level generic' STATUS d420c000b7080a13 MCGSTATUS 0 MCE 2 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC 1b2fb0be756ef ADDR 41391f60 Northbridge Chipkill ECC error Chipkill ECC syndrome = 2254 bit32 = err cpu0 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node origin, request didn't time out generic read mem transaction memory access, level generic' STATUS d42a400122080813 MCGSTATUS 0 MCE 3 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC 1b356485048aa ADDR 41393740 Northbridge Chipkill ECC error Chipkill ECC syndrome = b741 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node response, request didn't time out generic read mem transaction memory access, level generic' STATUS d420c000b7080a13 MCGSTATUS 0 MCE 4 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC 1b54c147b3d4f ADDR 41393fa0 Northbridge Chipkill ECC error Chipkill ECC syndrome = 3554 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node response, request didn't time out generic read mem transaction memory access, level generic' STATUS d42a400035080a13 MCGSTATUS 0 MCE 5 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC 1bc11900370aa ADDR 413918a0 Northbridge Chipkill ECC error Chipkill ECC syndrome = b741 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node response, request didn't time out generic read mem transaction memory access, level generic' STATUS d420c000b7080a13 MCGSTATUS 0 MCE 6 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC 1bc6ccc690629 ADDR 41393fa0 Northbridge Chipkill ECC error Chipkill ECC syndrome = 3554 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node response, request didn't time out generic read mem transaction memory access, level generic' STATUS d42a400035080a13 MCGSTATUS 0 MCE 7 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC 1bdac1fcc62df ADDR 41391a20 Northbridge Chipkill ECC error Chipkill ECC syndrome = b741 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node response, request didn't time out generic read mem transaction memory access, level generic' STATUS d420c000b7080a13 MCGSTATUS 0 MCE 8 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 0 4 northbridge TSC 1c197b828ca38 ADDR 413918a0 Northbridge Chipkill ECC error Chipkill ECC syndrome = b741 bit32 = err cpu0 bit46 = corrected ecc error bit62 = error overflow (multiple errors) bus error 'local node origin, request didn't time out generic read mem transaction memory access, level generic' STATUS d420c001b7080813 MCGSTATUS 0
Post 9 Posted by Jono 2008-12-29 08:47:14 UTC
Probably, looks like ECC is going nuts. Can you run memtest on the box at all? Pretty sure any new version of memtest has the ability to look for ECC events. Might need to let it run for a while, obviously a little hard on a live server :p
Post 10 Posted by Knight Of Nih 2008-12-29 08:50:08 UTC
could u run up memtest in a vm and it still tests the ram? i can imagine not but work investigating.
Post 11 Posted by DVD 2008-12-29 11:29:00 UTC
Post 12 Posted by Piggy 2008-12-30 08:17:36 UTC
Well i had a look around on the linux geek server forums
Non-fatal mce's are usually ecc faults, and *USUALLY* track back to bad memory, though it can also be overheating cpu, or a problematic cpu, or rarely the MB could be the fault.
ECC/MCE counts will get worse under load, unless the problem is really severe you won't see them at idle.
This is about all anyone had to say about it, most say memory is faulty, but can be other things as above
Post 13 Posted by Harbinger 2009-01-08 14:39:51 UTC
Server time off or is it just me?
seems to be nearly 40 minutes fast :S
Post 14 Posted by OverDrive 2009-01-08 21:02:55 UTC
Set to adelaide time and all fine - it has been crashing more often BUT has been auto restarting fine - just content to leave it atm until I can get the hardware done.