Ah ok, I’ve had these drives for a while now but ended up using SSDs instead so they just sat in my parts drawer reminding me of a bad purchase xd.
Crazy how they are like 3x the price now from when i bought them in 2023
Ah ok, I’ve had these drives for a while now but ended up using SSDs instead so they just sat in my parts drawer reminding me of a bad purchase xd.
Crazy how they are like 3x the price now from when i bought them in 2023
Do you happen to live in zuid holland?
I have 2 unopened Seagate IronWolf ST4000VN006 that I’ll be happy selling.
No idea what a fair price or your budget would be. That is if you are interested in these drives to begin with.
My 2 cents are that the issue is promotion not AI, if people started promoting stuff made without AI that would still be spam.
From the rules:
F/LOSS Exception: If your post is about a project that is completely open source & can be self-hosted in full without payment, your post is exempt from the 10% requirement. The exception does not exempt you from the account age requirement.
I would propose making this the requirement and not an exception, forbid all promotion of closed source, and allow the 10% requirement for open source projects.


Why would anyone pay for this when documentation is readily available, even for non technical users?
https://github.com/oobabooga/textgen
Installing something like textgen is as easy as downloading a single executable.


I’m running 2x4090, the 35B fits very comfortable in that.
For large models like the 397B without a ton of money there are several ways, ive seen posts of people using arrays of used 3090s with good results.
The other option is CPU inference although with current RAM prices that is less cost effective.
I was looking at maybe an array of Milk-V JUPITER2 since vllm added riscv support which could be very cost effective.


Depending what OP was using before but going from something like GPT5.2 to LLama 3 8B will be a massive difference (Although OP says to use it only for basic tasks so that does offset it)
LLama 3 already being a very old model doesn’t help either
I run Qwen3.5-35B-A3B-AWQ-4bit which while leagues ahead of LLama 3 8B still is a very noticeable difference.
This is not to say open source is bad, if one had the resources to run something like Qwen3.5-397B-A17B it would also be up there.


Haven’t used shotcut or openshot but can recommend kdenlive, it has alot of features but I find it a bit clunky sometimes to work with.
For me when I was working on a large project it crashed once and the autorecover worked, although I still manually save since I’m paranoid.


I’m selfhosting Forgejo and i don’t really see the benefit of migrating to a container, i can easily install and update it via the package manager so what benefit does containerization give?


Why do core counts and memory type matter when the table includes memory bandwith and tflop16?
The H200 has HBM and alot of tensor cores which is reflected in its high stats in the table and the amd gpus don’t have cuda cores.
I know a major deterioration is to be expected but how major? Even in extreme cases with only 10% efficiency of the total power then its still competitive against the H200 since you can get way more for the price, even if you can only use 10% of that.


Thanks! Ill go check it out.


My target model is Qwen/Qwen3-235B-A22B-FP8. Ideally its maxium context lenght of 131K but i’m willing to compromise. I find it hard to give an concrete t/s awnser, let’s put it around 50. At max load probably around 8 concurrent users, but these situations will be rare enough that oprimizing for single user is probably more worth it.
My current setup is already: Xeon w7-3465X 128gb DDR5 2x 4090
It gets nice enough peformance loading 32B models completely in vram, but i am skeptical that a simillar system can run a 671B at higher speeds then a snails space, i currently run vLLM because it has higher peformance with tensor parrelism then lama.cpp but i shall check out ik_lama.cpp.


While I would still say it’s excessive to respond with “😑” i was too quick in waving these issues away.
Another commenter explained that residential power physically does not suppply enough to match high end gpus is why even for selfhosters they could be worth it.


Thanks, While I still would like to know thr peformance scaling of a cheap cluster this does awnser the question, pay way more for high end cards like the H200 for greater efficiency, or pay less and have to deal with these issues.


FP8 Tensor Core, the RTX pro 6000 datasheet keeps it vague with only mentioning AI TOPS, which they define as Effective FP4 TOPS with sparsity, and they didn’t even bother writing a datasheet for he 5090 only saying 3352 AI TOPS, which i suppose is fp4 then. the AMD datasheets only list fp16 and int8 matrix, whether int8 matrix is equal to fp8 i don’t know. So FP16 was the common denominator for all the cards i could find without comparing apples with oranges.

?


Well a scam for selfhosters, for datacenters it’s different ofcourse.
Im looking to upgrade to my first dedicated built server coming from only SBCs so I’m not sure how much of a concern heat will be, but space and power shouldn’t be an issue. (Within reason ofcourse)
Do people seriously buy this?
Like i can only assume these are fake, has anyone ever seen this actually used?