lemm.ee

in asklemmy@lemmy.ml•How are we going to pay for all this?

8 points

1 year ago

Reddit has over 2,000 employees most of whom are doing bullshit nobody using the site actually needs or wants, it’s possible to run a lot leaner than that. Like Reddit itself used to, before they started burning hundreds of millions trying to compete with every other social media site at once instead of being Reddit

report

[ - ]

in localllama@sh.itjust.works•Discovering Locally Run Language Models: Share Your Favorites/Not So Favorites!

6 points

1 year ago

The wizard-vicuna family is my favorite, they successfully combine lucidity with creativity. Wizard-vicuna-30b is competitive with guanaco-65b in most cases while being subjectively more fun. I hope we get a 65b version, or a Falcon 40B one

I’ve been generally unimpressed with models advertised as good for storytelling or roleplay, they tend to be incoherent. It’s much easier to get wizard-vicuna to write fluent prose than it is to get one of those to stop mixing up characters or rules. I think there might be some sort of poison pill in the Pygmalion dataset, it’s the common factor in all the models that didn’t work well for me.

report

[ - ]

in localllama@sh.itjust.works•Discovering Locally Run Language Models: Share Your Favorites/Not So Favorites!

6 points

1 year ago

W-V is supposedly trained for “USER:/ASSISTANT:” but I’ve found it flexible and able to work with anything that’s consistent. For creative writing I’ll often do “USER:/STORY:”. More than two such tags also work, e.g. I did a rpg-style thing with three characters plus an omniscient narrator, by just describing each of them with their tag in the prompt, and it worked nearly flawlessly. Very impressive actually.

report

Why is the front page suddenly so stale?

posted 1 year ago

actually-a-cat@sh.itjust.worksOP

main@sh.itjust.works

11 commentshide report

[ - ]

9 points

1 year ago

in main@sh.itjust.works•Why is the front page suddenly so stale?

Top day is a good tip. Though I do think something is broken, seems too unlikely that this particular batch of shitposts is so uniquely hot it stays up all day when before the feed was moving quite fast

report

actually-a-cat@sh.itjust.worksOP

[ - ]

3 points

1 year ago

in main@sh.itjust.works•Why is the front page suddenly so stale?

all… Subscribed and local seem to be working correctly, some of the posts on them are new

report

[ - ]

in localllama@sh.itjust.works•Vicuna-33B-1-3-SuperHOT-8K-GPTQ

3 points

1 year ago

Unfortunately there’s just no way. KV cache size scales with the square of context length, so at 8k it’s 16 times larger than at 2k, for 33b that’s over 20GB for the cache alone, without weights or other buffers.

report

[ - ]

in localllama@sh.itjust.works•Vicuna-33B-1-3-SuperHOT-8K-GPTQ

3 points

1 year ago

That’s what llama.cpp and kobold.cpp do, the KV cache is the last thing that gets offloaded so you can offload weights and keep the cache in RAM. Although neither support SuperHOT right now.

MQA models like Falcon-40B or MPT are going to be better for large context lengths. They have a tiny KV cache so even blown up 16x it’s not a problem.

report

[ - ]

in localllama@sh.itjust.works•best method do use amd GPU for inference on linux

2 points

1 year ago

Deleted by creator

report

[ - ]

in localllama@sh.itjust.works•best method do use amd GPU for inference on linux

3 points

1 year ago

Not sure what happened to this comment… Anyway, ooba (text-generation-webui) works with AMD on Linux but ROCm is super jank at the best of times and 6700XT is not officially supported so it might be hopeless.

llama.cpp has some GPU acceleration support on AMD in CLBlast mode, if you aren’t already using it, might be worth trying.

report

[ - ]

in localllama@sh.itjust.works•best method do use amd GPU for inference on linux

5 points

1 year ago

I can recommend kobold, it’s a lot simpler to set up than ooba and usually runs faster too.