I just updated my setup from LMStudio to llama.cpp with the new QWEN 3.6 27B MTP model and I am getting 80-112 tokens/second, 90 average which is just shocking to me. I am on a 4090 with a context Window of 64k. It hardly use cloud AI anymore as I rarely need more than 64k if I ensure my first prompt is written like a design document. Multiple prompts are not great so I often just figure out where my initial prompt went wrong, adjust and try again in a fresh session. Way faster this way too. It has really worked out well for me as I am getting just as much done locally for free as I was with hundreds of dollar a month on cloud AI. I am still shocked and grateful it flowed this way.
ImmersiveMatthew
I am using llamma.cpp with QWEN 3.6 27B MTP, with a 64k context window on a 4090 that OpenCode talks to and then it in term talks to the Unity Game engine via MCP. Getting 80/112 tokens/second work 90 average which is shocking to me as it really does feel as fast as cloud AI (well faster for me as I am in Vietnam and round trips to US data centers really adds up in a session). The only really issue is you pretty much have to one shot prompts as follow up prompts will easily go over the context window size. If I cannot one shot prompts them use cloud AI both that is very rare for my use case. Maybe 1 in 50 or so and only when the tasks touches a lot of large scripts and scenes.
Agreed. I am not longer paying token fees as I am running QWEN 3.6 27B MTP on my 4090 GPU and it is as good and as fast as the frontier models for agentic coding.
The corporations certainly think they are setting themselves up for domination, but that is not how it is trending. The opposite is about to happen as all those over the top irrational exuberance driven data center investments are quickly becoming a liability as Open Source has caught up and Ternary 1.58 bit models are coming to make GPU redundant. Hate to be AI CEO or cloud ai investor right now.
I agree and would add for others reason here that the Luddites issue was not the looms they destroyed, but the out of control inequality that the government was not addressing. We need to stop blaming AI as a society for job loss and instead get the governments to help with the transition which so far they have largely been inactive on.
That was the part that shocked me as it rings even more true at this present moment.
Wow. I had to go verify this and sure enough, I found a video from 7 years ago where Lucas said this and much more including comparing the American Empire to the Galactic Empire. Wow. I missed that reference as a child.
Pathological Demand Avoidance is such a passive aggressive term. I find most in the spectrum happily follow a leader provided they are logical, compassionate, considerate and looking for win wins. If the leader is an egotistical boss who sees you as less worthy than them and is demanding you do work that makes no sense and is even perhaps harmful or exploitative, than yeah, I am going to avoidant and not compliant. Wish more were PDA as we would not end up with tyrant leaders.
It is not capitalism and this is not me defending it, but rather pointing out that it is how humanity naturally seems to structure itself. The narcissist / psychopaths / sociopaths tend to get into “leadership” positions and the dumb folks idolize and empower them and thus we end up with the smartest people writing a few paragraphs for the manager who is socially savvy but suffers from the Dunning Kruger Effect. This happens in all economic and political systems including Capitalism and Communism as the Sam people tend to get into positions of power.
Or cannot see it due to the Dunning Kruger Effect. It is very strong in some.
They are in on it or worse also caught up in the Epstein/Trump honey Trump.
The irrational exuberance is a drug and Jensen is in part one of the drug dealers who caused it.