JetBrains reports revenue growth, net financial loss for 2025
helgilibrary.com593 points by thw_9a83c 3 days ago
593 points by thw_9a83c 3 days ago
All the death knell comments - Is no one looking at the revenue line? Revenue still trending the same. Costs presumably haven't skyrocketed. They've invested in something big. Thats probably a good thing, and often needed to evolve.
Yeah look at the details below. Revenue is up 6% (which is not amazing but still growth). But staff expenses are up 34.2%. And the real big hint is "Cash Flow From Investing" plunging from -83 million $USD to -469. That's massive. They're staffing up and investing in... something.
> and investing in... something.
Reads like aimlessly burning whatever they felt they were able to burn without existential risk on some "also ran" LLM sinkhole.
In other words: nothing to see here, same as everybody else.
I still have this naive notion that we don't need LLMs for code generation and editing. Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
Maybe people smarter than me know better but couldn't there be a middle ground where an IDE/Editor has an embedded engine (doesn't need to be a full-on LLM) that doesn't require external tool calls and token spend?
If an organization is paying $2400/year per developer for tokens and a highly intelligent editor/IDE comes around that charges $1000/yr and gets more output at a fixed cost, its a no-brainer of a decision.
You know what, considering there was a recent "small" open weights LLM released recently that meets 90% of my coding needs I'm inclined to agree.
Qwen3.8-Flash-Next - relatively small, it runs on 6 6 year old GPUs on my home PC happily running 5 simultaneous 262k sessions with additional 10 cached in RAM (bought back when you didn't have to remortgage your house for Ram) and it has been the first local model that is not a toy.
But there is a class of problems where I still reach for Anthropic's fable...
However, I have a hunch bordering with certainty Anthropic is achieving such great results by doing a lot of harness tricks.
For example opus 4.8, is not much better on coding than before mentioned Qwen model, but gets amazing results on factual knowledge stuff (the knowing all works of Shakespeare thing). How hard would it be to add a general knowledge RAG to requests that contain relevant questions and beat all benchmarks like that? Not very hard.
So I think there is big innovation to be had in harnesses, routers, inference and so on.
As to money spent on AI per developer my current client (a fortune 200 software company) spends $500 per month. That is $6k a year. A lot more than your examples. And many people run out of their quota pretty quickly.
I was inclined to listen to your take… then you started talking about having 6 graphics cards in your computer and acting like that’s a normal thing that people do…
The other pitfall here is that we are tempted to compare this crazy setup to frontier models when the real comparison is running this same model in OpenRouter.
This 6 GPU setup will probably outspend OpenRouter on electricity alone.
These are 6 year old cards as one poster below pointed out. It is a normal thing people do, and many more people would if the prices were not made bonkers by few large companies vacuuming entire advanced chips manufacturing capacity just to lock those chips in warehouses.
If all those nvidia gpus the hyperscalers bought were online the cost to rent a single b200 wouldn't be $50 and hour and you could buy those 6 year old gpus I use for $200 each. Not almost $2k a pop they sold for now.
> the cost to rent a single b200
A 180gb b200 on runpod is $6.8/hr. Nowhere near $50/hr.
The _most expensive_ on-demand b200 rental I can find is $11.2/hr. There's a number of websites around these days that track GPU rental prices across different services: gpus.io, gpu.watchworks.dev (vast and runpod only), gputracker.net (not free, lol), gpurentalprices.com, priceofcompute.com, rentgpu.org, etc (I'm only looking at the first page of search results for GPU rental prices tracker search). Most of these don't include vast or runpod, but all of them show prices under $10, mostly around $6/hr for b200.
no typical person is going to buy 6 GPUs unless those are 6 dies on a single card.
Ergo 6 GPUs isn't a typical setup, never has been and most likely never will.
OP is a classic outlier dev who is able and willing to set up and maintain such a thing - most normal people aren't.
Yeah but they're from 2020. Normal people in a few years might have that in their laptops, just clocked way down to save on power.
That's what, the same as a Mac studio?
Keep in mind that GPU iteration cycles are slowing down significantly for some time now, so "a GPU from 2020" can actually mean NVIDIA RTX 3090 with 24 GB VRAM, one of the best cards you can buy in 2026 for local LLM use; in fact buying one used is probably the best choice right now, because newer generation RTXes are crazy expensive.
My GPUS are Rtx3090s with 24GB ram. You guessed correctly.
I do not enjoy the fact those cards cost more than their msrp 6 years later, but there is a much more important consideration than money (which also makes sense, but about that later). It is the fact soon people will not be able to do my job without AI at all. Even now if I didn't use it I think I'd be out competed very quickly.
And having the ability to run it locally, using a really useful, not toy model is very useful. It makes you independent from Anthropic deciding to ban your account for example.
As for money, it is an open secret the biggest cost of coding agents use is input tokens not generation. I tend to use about 1.3B input tokens per week on claude code with only 7-8M out. Out if this 80% is cached. And the cache is pretty restrictive. You have 5min cache and 1h cache. If you don't keep reading over that time your cache expires on the cloud. Then your 500k context counts as 500k input in its entirety.
And the numbers I mentioned would cost thousands of USD a week at API prices.
But when you control inference, you can keep your cache for as long as you want and save it to disk.
I tend to have up to 10 coding agent sessions open at a time. Some are used once a week. I never use more than 5 at the same moment. Having 15 full contexts cached in RAM basically moves my local cache utilisation to 95%+
Basically I think the AI companies will soon require us to pay the real price for the inference. I prefer to be ready.
I have two. I paid <$1000 for each, before this craziness started. They work, sure, but 48GB of VRAM is just not enough to use a local LLM like a cloud model.
we use that same model at our company, it powers not only our devs but also many business needs. We rent 1 gpu (B300), it costs around 10x less than the api costs
Do you rent from AWS or some other provider? I wonder if there are issues with on-demand rent for those high-end GPU instances, is capacity always there or sometimes it is unavailable?
I don't touch AWS. We rent from an european provider. Demand can be tricky for serverless/spot-pricing. We have an always-on server so, no issues.
If there's some cost or availability question you have, AWS has some complicated bullshit to solve the problem for you. Check out reserved instances.
Can you tell more about the class of problems that need Fable?
Not the person to whom you responded, but several times I have been dealing with debugging. Issue, tried. Opus/Sol/whatever and come up dry, then thrown it at Fable and gotten the solution in one shot.
One example of a "hobby grade" problem that only fable could resolve that I had no time to mess around with.
I have a home network consisting of multiple buildings, a server room with a k8 cluster and various devices, some Cctv cameras, redundant fiber links between buildings, ftth Internet and lte backup, and so on. Not a simple network. All on Mikrotik switches using a lot of modern features like L3 in hardware routing. All properly designed, servers are multi homed with 10G DAC cables between switched.
Occasionally I'd notice few second drops when observing Cctv from my cameras on the monitor attached to my pc. Pc on the 10G.
Also occasionally I'd get random devices (android TV) decide "it has no Internet) for a minute at a time.
Also occasionally I'd have my AI take 20s to answer when I know for a fact it is doing absolutely nothing.
First I looked into it myself and I found nothing. No misconfiguration etc. Then I used opus with it. It found no network drops on any interfaces etc. But it wrote a script that opened 10k connections, kept them open and sent traffic through them. This was tried to all my k8 nodes and one node would occasionally refuse to open 0.1% of these connections.
Opus decided it must be a network card or a DAC cable. It took ages to identify which. I replaced both, problem went away for weeks each time and then returned every time.
Opus had a bajillion ideas to mess with my network config. Thankfully I know enough about networking not to let it take me on a wild goose chase.
Then I used Fable. Fable took 5 minutes and found the server grade Nic card's driver I use has a known rare problem where in certain configuration the default memory buffers for some hardware offload feature are too small for it and when connections get opened rapidly it sometimes chokes.
But there are two nodes with exact same config. Why only one was affected? Actually both were affected, but on one it was so rare I had to run the testing for hours to notice it.
The memory was increased and the problem... Became a lot rarer. Not resolved completely.
Fable again. This time it came up with an idea there must be a bug in the active/standby part of the driver where it occasionally let's some traffic through the standby interface. Which makes the switches ARP learn the standby as the correct path to that mac and send a portion of that traffic there, but the standby couldn't receive traffic.
It setup tcpdump on both the standby and live and proved indeed it was happening. I do not remember how this was resolved, but it was and for last 4 weeks the problem is gone.
I have actual programming problems too. 4 of which I turned into a personal Ai benchmark. Only fable solves all 4.
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
Turns out that actually - no. Researchers have managed to prune half the Experts in a MoE model that had a low probability of getting activated during coding tasks, resulting in a more focused model:
https://arxiv.org/abs/2607.16721
Main benefit is that it greatly reduces the amount of RAM required to run these models. Of course you could just cache those unused experts on disk instead, but the main point here is that you know which ones matter.
But aside from that recent models, like Qwen3.8-27b are reportedly more durable under heavy quantisation, e.g. 3bits or even ternary. With additional techniques like TurboQuant, you can feasibly run these models on consumer hardware - even if at 1/4th the speed you'd get from rented infrastructure.
VS Code has extensions such as Kilo Code or llama-vscode which let you work with local models much like you would with cloud based solutions.
> "Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?"
If the implementation brief says "attempting a reconnect in this handler would be a wild goose chase", the model needs to know enough Shakespeare, at least indirectly, to understand that expression...
Forgive my naive understanding of LLMs - but how do you get semantic understanding of a codebase, such that it knows what changes to make/why/where, without a wider understanding of language more broadly?
I'm using the word understanding loosely there, but I couldn't think of another word.
Depends on what exactly you want to change. Lsp can do a lot but only in very simple changes.
Intellij when I used to use it had a lot great features like refactoring, extracting part of code as a function, renaming and creating empty classes/boilerplate but that's it
In current job I can order LLM to take data sink from other endpoint and write new with given URL. It will fetch from endpoint, check what it gives, compare with other and write new sink. Then it needs polishing because it always create something as awful as possible with cloning data all around but the most boring and soul sucking part is done
This is why I am working on https://github.com/spockz/semantic-editor. To bring more powerful editing functionality to agents. In a way that attaches to their chain of thought and deals with their probabilistic framing in json so it all works out cheaper and faster.
LLMs at my work have access to Slack chat history, the Wiki, the complete git history, (AI generated) notes from every in person meeting team meeting, JIRA, all code reviews, and all production logs, as well as your personal email (when you run it yourself)
They can typically explain the history of something faster than any human
How does your agent wire all the information? Does it use MCP for all of them or do you use another method to access or collect the data? and how does it resolve contradictions (one source say one thing, another say otherwise).
It's a totally fair question. I'm personally wondering if there's a half way point. Some kind of structured language that isn't plain English that a "dumb" LLM is able to parse. It could be human written, or it could be written by a "smart" LLM at a greater cost.
The more we go in that direction we closer we get to reinventing frameworks in a more compute heavy way.
It just so happens that there's no framework (outside silos) for many specific tasks, and AI is the workaround to surface those patterns.
Looking at what Jev has shown, and what is being done in the space. I think you are right that there’s probably a lot of room for performance improvement on specific tasks and workflows that will be happening in the next few years
I also imagine it could be a big shakeup if all of a sudden models could run on CPU. Imagine running an Astra-level coding agent, locally on your laptop. All of a sudden GPUs wouldn’t look as valuable, if you don’t need them as much
We are still some time away from that, but it seems like progress is being made
We’re conditioned to interact with language models as chatbots and in that sense strong language understanding (implicit - some sort of world knowledge), is probably necessary for that?
But I’m sure we can have a small model that’s really strong at programming concepts, JavaScript syntax, and that’s about it. You’d interact with it differently, at specific seams in your code base - review a PR, merge two functions together, investigate these logs.
Or maybe I’m just not adequately absorbing the bitter lesson. Idk
I think the amount of knowledge to correctly work on code is more than you'd think, because at the end of the day writing code without an understanding of the environment it exists in/for is likely to not fit the problem correctly. Maybe it doesn't need knowledge of _Shakespeare_ per se, but if you were working on a virtual tabletop having knowledge of tabletop games can help with identifying the right implementation to use, knowing what kind of constraints to consider, etc.
> If an organization is paying $2400/year per developer for tokens
My employer is spending $50,000/yr per employee on tokens, and they're not alone
I would say paying 50,000 USD per year for tokens to double productivity is cheap.
And hopefully they are just paying that per developer, not per employee.
No kidding. As a small employer I expect to be spending $200 a month so each employee can have an OAI or Claude 20x sub, or else spend it on the Chinese prepaid tokens of their choice.
I have a feeling that when all the AI-hype dust settles, what you describe will be the killer app of AI. That, ChatBots and unstructured data processing. Huge productivity improvements but not the sci-fi hype of today.
I'm not an expert, but I think that:
1. Storing Shakespeare's work costs almost no $ in regards to disk space.
2. If the prompt doesn't include "Shakespeare" or relevant terms then no regression is performed for that topic and therefore there is no effective token cost.
Someone may correct me, but I think it's not a big $ win to exclude relevant topics from the models' overall capabilities. Instead you'd tune weights so that #2 better identifies what is or isn't among the relevant terms on which to run regressions.
I think you're probably missing the forest for the trees here... broadly speaking, these multibillion "parameter" (whatever that actually means) models store a lot more than shakespeare, and have storage costs in the 100s of GB/TB (which translates to $$$$$$ in SSD/RAM costs), nevermind the (kilos/mega/giga)watts involved, all the pollution, etc...
Meanwhile, a template (maybe a couple KB) costs less than a couple cents to store and run. Large Languages Models are not really interesting, (smaller) LLMs that only contain "what you need" are.
To me it feels like the difference between crows and humans. Yes crows are smart, but what we really want are all inclusive models capable of human "thinking" . I guess we probably need a mix of both so we can give the crows the easy jobs freeing up resources.
Right but if the cost difference is negligible, as I've pointed out in another comment, in what case would you prefer to tell the crow twice to do something that a better weighted large model could do in one prompt?
I don't see it. In every case my time is more valuable than the cost of the prompt (so far) so the higher dollar cost, one shot, "getter done" model is the better net value option.
This is especially true when taking into account that the crow doesn't just fail on a single prompt, it does something much worse. It creates new problems that need to be undone afterwards. It confidently duplicates, triplicates, etc. a damaging work output that then needs more and more work to clean up before starting over.
Claude Fable 5.1 currently estimates its own total data set at 45TB.
That's a trivial number for sharing between a small user base of, in my case, 150 employees.
What am I missing? You're talking about maybe $60k retail cost in high read speed SANs that are likely already in place for a business of this size anyway? (Probably purchased a few years ago for under $30k. At least mine are.) In a US data center I'm paying a flat rate for rack space, so power consumption isn't a consideration anyway, but honestly it really isn't that much power even if I was being billed for it.
It's really nothing. An overlooked line item on a budget sheet.
So the real cost if I were to self host is video cards. But as mentioned previously, that's not reduced by smaller data sets, it's reduced by better weighted models, right? Let me know if I'm mistaken, please.
I'm very interested to be persuaded otherwise as a decision maker. Thanks for your insight and ideas.
Considering all of the above, I'm currently of the mind that a lower cost model that makes more mistakes is much more expensive in net, actually. So I prefer the most accurate, better weighted model, not the slimmer data set.
> Maybe people smarter than me know better but couldn't there be a middle ground where an IDE/Editor has an embedded engine (doesn't need to be a full-on LLM) that doesn't require external tool calls and token spend?
IDEA already have small LLM for one line code completion IIRC.
But the gain people want from LLM is generally "here, add this entire feature" or "here, go thru every dependency's changelog and update code to work with latest version". Those are not small LLM tasks
Great idea! It’d be something like Common Business Oriented Language, or COBOL for short.
I think this already happens through Mixture of Experts which is now build in to ost models.
But finding out what an LLM needs to understand from a business side to write your code good, is an otpimzation which no one cares currently.
I'm pretty sure we either stay on big full frontier models for a long time, just use them for everything or we will start to see more and more people doing finetuning/project specific training like java + german + english + business contxt xy;
It will be an indicator for the whole industry.
My gut says the economics, e.g hardware/data center/resource constraints, are going make the economics of small specialized models more attractive. Without any evidence whatsoever, I also think that the big frontier companies will have to de-emphasize chatbots as huge models in favor of chatbots as huge products with a much much more granular mixture of experts approach, but with much smaller models. I’ve been saying for a while now that AI products have to hit the gas on prioritizing product design to reliably solve real people’s problems in predictable-enough ways, because the current approach is only really appealing to enthusiasts, developers, or optimistic managers, and with the kind of money they’re throwing around, that’s not going to work.
My gut is that you could distill 98% of what's currently being done inside LLMs into a classical knowledge base/inference engine and only use the language models for natural language processing and an oracle for brainstorming and cut most of the computational cost and hallucinations out.
I don't need AI in the same way that I don't need autocomplete. I can definitely program without autocompletion, but I'm a lot slower than others who use it
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
The same can be said about software engineers. The tricky thing is, it's hard to separate a subset of knowledge from the whole, for both human beings and LLMs
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
Probably not unless you're writing tooling relevant to literature or prose, but I can't imagine trusting jetbrains (or any ai studio) to curate this.
It is very difficult to predict ahead of time what knowledge a model would need to understand a prompt to generate a program. It could refer to all kinds of real world knowledge referring to the kinds of entities you want the program to model.
I have hope we will get there eventually, once all the hype/wealth extraction/boys club giving all their buddies money cycles end, and the specialized tools with real value start to emerge.
These specialized tools already deliver tremendous value. What happens on the backend financially is of no concern to me, as I have no influence over it.
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
What language do you plan on prompting it in?
I plan to use english, but not 1500's english.
That might be a bad example ... replace english with nearly anything not related to prompting/coding. For example, I bet the models have "knowledge" of biology, chemistry, etc. not exactly useful for programming a SaaS web app that say does project mgmt. I think there's opportunity for very specific tooling rather than "general" knowledge.
I don't train LLMs myself, but my impression is that general knowledge has lead to more intelligent behavior on specific tasks. The voluminous training set is imporant.
But even besides that... Yesterday I prompted a feature by referencing a specific Monty Python skit. Does the coding tool need to know Monty Python? Maybe I could describe the feature in other terms, but it sure was convenient to have this shared knowledge. I don't see why Shakespeare would be any different.
CodeRabbit emits LLM generated poetry in its comments so maybe more than you'd think?
I suppose how much literary support you need in your model depends directly on how erudite the comments are.
Would there be problem domains in which the more educated LLM would perform better? Are your names directly related to concepts from said domain,
LLM comments: "I think it may be a potential bug that the sum VATAddedTax gets added to the TaxFreeItems".
I mean it seems a bit unnecessary but also maybe it can help in unexpected ways.
Because we don't need code
func randomName () { desired machine physics }
Everything around "desired machine physics" is superfluous wank; historically a biz case stored as code when some UI could feed biz case params go a function generator
Come on we know what we use computers for; media consumption and 2D data entry/review. Locally we just need a core engine for geometric transforms of visual state. What all these languages give us ability to create such a generic VM filled with customized semantics that mean nothing to solving the problem but plenty to a clever coder.
Kind of like Unicode we need distilled geometry primitives like "teapot for text" and desktop metaphors and to let people put the superfluous wank at the presentation layer
Which text used to be so making UI out of layers of text, OOP, and such made sense for decades
But we're just engaged in bloating system state through def jargon_to_encapsulate { desired machine physics } when we already know it's going to be simulated 3D or 2D visual transforms. We don't need to capture all those states in code verbatim.
Things like Jev are the future of models. Fine tuned on transforms given a context. "So you want to replicate GTA5? Here's a data set of geometric shapes and gradients constraints from all observed xyz" pipe that into your local renderer
We're entering the phase of software engineering (and engineering generally) where we realized we been dramatically over playing the song and can strip out entire asides and digressions, circumlocutions of provenance, to tighten up pacing and improve enjoyment of the outcomes. Hopefully. Or we kill ourselves. Through social squabbles (political, economic, religious, whatever) due to laziness to learn etiquette, and environmental destruction.
Did you learn english or C first?
I learned C first, then Perl. Sometimes I get an epiphany that terms I use in programming have meanings other than what it means in programming.
But you learned English (or some other human language) before either of those..
Well, the post I’m replying to says that English was learned first, and I’m saying that I learned Perl before I learned English.
I don't think LLM needs to know Shakespeare to understand: "attempting a reconnect in this handler would be a wild goose chase." I would even bet that most people learned of this expression before reading Romeo and Julie.
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
That's needed to interpret the 10000 monkeys typing requirements in the various corporate product roles /s
> without existential risk
To do nothing is to die. I've been a JetBrains subscriber for over a decade. Nobody is going to use their tools in five years. I certainly don't need my subscription anymore.
> on some "also ran" LLM sinkhole.
They can't do Junie. It's a dead end. They can't do the same thing everyone else is doing either, or they're exactly as you put it: an "also ran" in a very crowded field.
The only way for JetBrains to survive is to figure out their Garmin play. They either find some niche within the existing product or - maybe (and very improbably) - they can innovate something wild nobody else has figured out yet that provides a path to new fertile pasture. But that's more the startup path than the thriving incumbent facing innovator's dilemma path.
But no matter what, if they stay the course, they're dead. Just like half the people expecting to still be writing code by hand.
I think they are trying new things even with "nobody writes code anymore" in mind. Their IDEs are now MCP servers for agents to reason about the code. Imagine your agent will stop wasting tokens using grep to find symbol usages if he can just ask IDE where it is used. They are not the best option for vibecoders right now but who knows what will happen in 1 year.
Integration with other developer tools, code reviews etc. are better than ever. Ideavim is probably the best vim plugin and it gets new features every month. They are also investing a lot in the AI.
I would bet they will survive just fine unless we would write software in slack only using emoji :-)
Imagine your agent will stop wasting tokens using grep to find symbol usages if he can just ask IDE where it is used. They are not the best option for vibecoders right now but who knows what will happen in 1 year.
Can't agents just use a language server? There are already a bunch of projects that provide this, e.g.: https://www.agent-lsp.com/
Doesn't seem like a huge differentiator anymore?
I don't really know much about their current users, but from a business perspective, post-tokenmaxxing and them having an IDE, it seems like the most efficient environment for having a human in the loop is something they could tackle?
Well nobody needs to use IDE. There were always people who used vim with tmux as their IDE. The value of IDE is the I as in integrated - I see some value in having easy to manage single dev environment with batteries included. (I'm not an IDE guy though, and I personally closer to vim/vscode + plugins)
There were always people who used vim with tmux as their IDE.
But until language servers, IDEs with built-in code analysis had a huge advantage. This stopped once LSP was designed and some language servers got mature. LSP was developed originally by VS Code and it is easy to see why - it allowed VS Code to compete with IDEs without writing their own code analysis by letting language developers write a standard server. This made VS Code (and others like Zed) so powerful, that they had already replaced IDEs for many users (maybe outside Java). LLMs are just another nail in the coffin.
> This stopped once LSP was designed
No, it didn't stop. In my experience, LSPs are still incredibly poor and comparatively primitive vs the actual analysis that JetBrains have implemented. It's night-and-day different in terms of the supported refactorings, actual understanding of the types and the codebase etc.
Also, LSP is just poorly designed and inefficient. JSON-RPC was an awful choice, the specification text is not well-written and the whole thing feels like a Microsoft pet project for VS Code than an open specification that everyone agrees works well.
I think the name misleads people about LSPs. LSPs are basically plugins for VS Code. The protocol semantics are visual, so they're tied to what VS Code wanted to display on screen. They don't really change the game, especially as you always had the option of using IDEA's own plugins in other tools and text editors (yes! there is an API and you can start up a headless IDEA to access it, even from programs written in C or Rust).
Not many did that because JetBrains never cared much to properly document this path, and so using LSPs is a better paved cowpath. But that's because IDEs aren't a real business for Microsoft, whereas they are for JetBrains, and we know where that leads - the landscape is filled with the skeletons of dead IDEs that were just loss making corporate side projects, defunded and "donated" to some foundation once the executive sponsors moved on. NetBeans and Eclipse are two of the most obvious but there have been others.
The risk with VS Code is it goes the same way. Eventually Microsoft needs to cut back, perhaps due to needing more capital for AI or due to AI related losses, and in the general layoffs that follow VS Code gets cut back to a skeleton crew.
This is not the whole picture. IDE is not just for writing code. What I really like on jetbrains is how strong the platform is. I will show you example which would be great if they implement it.
If you use ideavim you can enable feature which shows you action ID you have currently triggered (Ideavim: Track action IDs). If you click on anything it will show you name of that action and you can setup vim shortcut for that. This way, you can use almost anything by just typing a few letters. But that is just a start.
Now imagine you can use this to create "macro" which will create something like skill for LLM. So in a few seconds to minutes you can create "skills"/runbooks for major refactorings, project updates, git bisections, project deployments, log analysis or whatever you can think of. You will just click on 20 features and it will record what you did with some additional context (files, connected services, build systems etc.). This is the strength of integrated tools.
Of course this is just my imagination but I am not the smartest person in the world and I am pretty sure/hope that someone in jetbrains or in other IDE company is thinking this way.