Curious about the quant tho.
Q8 from unsloth.
Something like Qwen3.5-122b
My go to model for knowledge. Definitely much faster at Q5 but it lacks the tool calling quality of the Qwen3.6 models. Really hoping we see a Qwen3.6-122b soon…
Curious about the quant tho.
Q8 from unsloth.
Something like Qwen3.5-122b
My go to model for knowledge. Definitely much faster at Q5 but it lacks the tool calling quality of the Qwen3.6 models. Really hoping we see a Qwen3.6-122b soon…
About 200 t/s prompt processing and 10-20 t/s with MTP.
Greatly depends on the task, predictable things like code generates at 18-20 t/s. Creative writing more like 10-17 t/s.
Yes, I got a Strix Halo machine before the RAM price hike and use it to run all my ML stuff on it.
Currently using llama-swap with llama.cpp/ComfyUI and opencode/Open WebUI as frontend.
I’m running Qwen3.6-27b, Voxtral Mini 4b, Piper and Qwen Image. Also, some embedding and reranking models.
I use them for:
So excited for Bigscreen. Might just finally replace my Kodi setup.


It does? Guess I can finally yeet Chromium from my machine then.
If you have trouble with outgoing mails, you can use a hybrid approach.
Receive mails directly to your server but use a mail service to relay your outgoing mails. Configuration for that is very simple in mailcow and there are a few dozen (free) transactional email providers (e.g. Scaleway).
That way you can keep receiving your mails privately and only have to give up some privacy when sending mails.


Plasma Bigscreen is scheduled for its first release next month: https://plasma-bigscreen.org/


As a side note, Qwen3.6-27B is much more capable than Qwen3.6-35B, even though it is much slower.
https://huggingface.co/unsloth/Qwen3.6-27B-GGUF
For coding tasks where you don’t mind waiting, you should be able to barely squeeze in the 8-bit quantized version with 32 GB RAM + 8 GB VRAM and have a pretty competent local model. 4-bit quants work but they have issues with complex tool calls.
If you use the MTP branch of llama.cpp (and a suitable model) you can even double or triple your token generation speed: https://github.com/ggml-org/llama.cpp/pull/22673
For easier tasks, disable reasoning for instant responses.


That looks pretty good. Looks like Portainer is getting replaced this weekend.
Do you actually train the LLM or use RAG? I have been looking for a local LLM + Wikipedia RAG solution for a while now.
For now I just have kiwix-serve + searxng doing a simple search but the Kiwix search is…questionable.


I wrote an application which runs on my server and monitors my favorites on Tidal/Deezer/Qobuz. It downloads them in bulk whenever I have a premium account with one of them. Usually I purchase a month of premium every few months, at which point I get nice clean FLACs for local use.
The FLACs are moved to Jellyfin and I stream them using Finamp, which also supports transcoding, so I keep 128 kbps Opus files for offline playback and stream the raw FLAC files when bandwidth is no concern.
I have amassed a huge music library over the last decades, so even if all streaming websites go under tomorrow, I have enough music locally to last me a lifetime.
I am happy if someone uses AI first to come up with a coherent message, bug report, or question.
LLMs do not add anything of value to bug reports, they add unecessary padding requiring me to filter out the marketing speech to get down to the issue. I would much rather have the raw brain dump of theirs.
If somebody sends me their ChatGPT text I now ask them to send me their prompt instead so I don’t have to waste my time on their lengthy text that has the same amount of information as the original.
I am annoyed if it’s ill-researched/understood nonsense, AI assisted or not.
Being coherent is rarely the problem in bug reports, it’s the user not properly typing out what the actual issue is.
I have gotten bullet point list bug reports that read like they were written by an insane person that were more useful than a nicely written ChatGPT message with 0 information in it.


If we’re talking about online editing, Collabora has web editors based on LibreOffice but with a modern UI: https://www.collaboraonline.com/
They are really great and can be self hosted (e.g. with Nextcloud).
For offline editing, as already mentioned, LibreOffice has an optional ribbon UI and OnlyOffice looks pretty modern as well.


Yes, that’s still a bit annoying unfortunately.
Editing the fstab to properly mount a network share also currently has no UI available in KDE and has to be done manually.


but to discover it on my other linux machine is always a chore that involves editing a few config files and just kinda randomly poking around until it works.
What’s your desktop environment? On KDE you can just enter smb://serverhost/path in the Dolphin navigation bar and it will open it.


Indeed. Connections to my Tor bridge dropped by 80% when Iran disconnected.


The first one I think is a fundamental limitation in that display preferences by default is per-user. Maybe this makes it work for you? https://feddit.online/post/1350756/comment/6636228
I don’t really have this problem anymore since I got rid of my projector which advertised a resolution it couldn’t handle. Had to login into the void since the login screen never showed up. Looks like this might be fixable nowadays.
The 24h clock might be similar - check your system-wide locale.
The locale is set to American English but the time format is set to German, something the lock screen can handle but SDDM cannot. I also tried applying the Plasma settings to SDDM a few times but it doesn’t really change anything.


It always chooses the default highest resolution, (which may not work on some devices with faulty EDID), does not respect the Wayland/X11 choice, has a long pause when going from login screen to desktop and does not support 24h clocks.
Just to name a few.


How drunk are these guys?
Ask the dude that renamed Twitter to X (formerly Twitter).
What did you flash on your Kobo?