☆ Yσɠƚԋσʂ ☆
- 1.1K Posts
- 814 Comments
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•Kalashnikov Group presents Kalitka-KA anti-drone system
5·18 hours agopretty much the same thing, but with a human operator
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•An excellent blogpost from Shengyu Liu, kernel engineer at DeepSeek
3·1 day agooh haha didn’t proof read auto translate
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
3·2 days agoI don’t disagree with any of that. But I think we’re talking about different things here. My point is that it’s not clear that capability will continue to scale in a useful way just because you make the model bigger. If you keep getting diminishing returns while needing vastly more resources, then it’s not economically viable to run these huge models.
So, I expect that labs focusing on more efficient architectures will outcompete those that are trying to brute force the problem. Like sure, DeepSeek isn’t small in a sense that you can run it locally, but it is small compared to other models in its class, and much more energy efficient. Whatever hardware we get down the road is going to benefit more efficient models the same way meaning that they will always have a competitive advantage.
From what I see in the latest releases from Anthropic, Fable isn’t a huge leap ahead from Opus. There is an improvement, but it’s not a definitive jump in capability the way it was from Sonnet to Opus. So, they managed to make a bigger model, but got diminishing returns, and it’s evidently so expensive to run right now that they can’t even offer it as a default.
The real progress will almost certainly be happening in hybrid architectures where people start coming up with algorithms that complement LLMs and augment their capabilities. These will be like different brain regions responsible for different tasks. For example, memory formation is an obvious example here, another would be to have a built in mathematics engine. A real huge win would be to figure out how to do few shot learning on the fly as well, for which memory is a prerequisite. So, there are plenty of things we already know that can be done much better.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
2·2 days agoThe problem is with the context and data propagation through the network. As you keep making it bigger it becomes slower and less focused. And there is research showing that smaller models do outperform large ones on some tasks https://cacm.acm.org/news/bigger-not-necessarily-better
What I expect we’ll see going forward is more hierarchical architecture where you have finely tuned models for specific tasks with a general routing model on top. This is basically already where MoE architecture is moving now. We might also see stuff like neurosymbolics get more popular where the LLM acts as a stochastic engine within a symbolic logic system. The model can handle noisy input from the real world, and transform it into structured data that a symbolic engine can operate on.
Brute forcing the problem is a naive approach and US labs took it because they effectively had unlimited resources to train their models until now.
And when more compute becomes available, solutions that are more efficient are going to further benefit from that as well. We see this with DeepSeek right now. They focused on efficiency over capability up front, and now they have a fundamentally cheaper architecture that’s rapidly catching up in capability.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
71·2 days agoAnd the big problem for them is that they have no leverage of Chinese labs, and it’s hard to ban open models.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
2·3 days agoThat’s definitely a plausible option, but it’s going to be very hard to ban use of open models. They could get use of official Chinese services banned, but justifying why OpenRouter and others can’t run them is going to be a lot harder. And there’s also a ton of money invested in all these AI companies running on open models now. So, the pushback will be significant.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
5·3 days agoThe problem for them could end up being that the economics simply don’t work. If more capable models are more power hungry, then operating them might be too expensive to justify. Or it could be that there are diminishing returns, and they simply can’t make a model that’s significantly better than the current frontier.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
7·3 days agoI’m hoping Alibaba will start selling these things at rpi prices https://wccftech.com/alibabas-tsmc-built-5nm-risc-v-chip-xuantie-c950-now-runs-qwen-3-8-27b-model-natively-unlocking-massive-vertical-integration-tailwinds/
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
2·3 days agoAgain, there is no reason to think that you can just keep making the model bigger and keep getting improved capability that way. In fact, we already know that’s not the case because simply making them bigger stopped being the focus. The real breakthrough is going to come from better algorithms.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
8·3 days agoThey have no leverage over Chinese labs, and China has every incentive to continue developing this tech. The only real explanation I see here is that they’re starting to get into diminishing returns territory, investors are getting edgy, and the costs of running this stuff are exploding.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
10·3 days agoOh they definitely aren’t, there’s an interview with Alibaba Cloud founder where he discusses the direction in China. Basically, the goal is to find useful niches for this tech early on, then iterate and improve. They’re not chasing AGI or trying to make one model to rule them all. That said thoough, the capabilities of Chinese models in the same domains where American ones shine are very close as well. So, I do expect that Chinese models will catch up and start surpassing American ones on their own turf before long. I’m also expecting that the trend will shift towards running smaller and local models for most things because you just don’t need a giant model to do most tasks.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
9·3 days agoI don’t see how anything they do can possibly affect what Chinese labs are doing. And that’s the only alternative to American labs right now. So, who are they going to convince exactly?
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
9·3 days agoI meant that simply making models bigger might not actually make them more capable. So even if you had unlimited hardware to play with, you might have to find a different approach.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
4·3 days agoI expect so as well, and my prediction is that we’ll have LLMs that are roughly as capable as the current frontier that can be run locally within a year or two. At that point, it’s just going to be good enough for vast majority of tasks most people need to do.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
15·3 days agoThere’s no reason to think that the architecture itself can scale indefinitely. It might very well be that LLMs have some hard constraints on the scope of the problems they’re capable of solving.

















they’re not https://globalnews.ca/news/11840683/ai-china-layoffs-court-ruling-canada/