Rendered at 23:38:28 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
sicktriple 7 hours ago [-]
Man, just when it seemed like we had it all. Computers were cheap, efficient and powerful. Somehow we figured out a way to accomplish tasks we already had solved except now it's 1000x more expensive, requires the combined electricity of the entire world, is reliant on someone else's rented compute, and now I need a robot to tell me what button to press. What a time to be alive.
scotty79 2 hours ago [-]
In 40 year software designers couldn't figure out that there should be a search function for every functionality that your computer provides.
When we got a form of such search, it searched all of the internet alongside with the functions of your software on your computer. You had to tell it manually what software you have and when it found something you had to fish out the description of what to click to get the functionality you are seeking and then click through some dumb ui to actually invoke it.
Agents are the first thing that can lead you directly from "I know what I want my computer to do." to "Actually doing it."
Success of agents is founded on the profound and sustained failure of all of the software designers and developers ever.
1718627440 2 hours ago [-]
> In 40 year software designers couldn't figure out that there should be a search function for every functionality that your computer provides.
apropos ?
alanbernstein 6 hours ago [-]
Valid points, but this sounds extremely useful for learning complicated GUI applications like 3d modeling or media editing apps.
ASalazarMX 6 hours ago [-]
Valid point, but the solution is wasteful and overengineered. Imagine if videogames needed an external datacenter to run the tutorial levels.
alanbernstein 1 hours ago [-]
I don't disagree. But the existing tech solution - scrubbing through YouTube tutorial videos until I cobble together the bits of knowledge I need - is also quite overengineered, isn't it?
Obviously I'd prefer to ask specific questions to my expert human friend sitting at the next desk over, but sadly they don't exist.
I just hope there is a future where this tech exists without such big downsides.
conorcleary 56 minutes ago [-]
There'll be gauntlets of pathways mapped before the first step is taken, as or before the first agent is deployed. Criteria will include resource use weighted against economic efficiency. Already being done.
arcanemachiner 4 hours ago [-]
I am willing to bet you could run a model capable of using this on a 12GB 3060 + some RAM. (Qwen 3.6 35B)
Even the data centre using this probably uses less net energy, and costs less, to build one of these dumb arrows than the aggregate sum of the energy used to keep you alive while you Googled for the answer. (Including heating/cooling the building, powering the equipment growing the food you eat, etc.)
fasterik 5 hours ago [-]
Sounds like an opportunity for a good engineer to come in and write something with the same functionality that runs locally and efficiently. The claim that it's wasteful only holds water if the non-wasteful solution exists and can accomplish the same tasks.
alanbernstein 1 hours ago [-]
Right, maybe the next step is that Blender comes with a local-gui-tutorial-llm packaged with it?
ramon156 4 minutes ago [-]
isnt the LLM the wasteful part?
Nonetheless why not let an LLM make a guide and use that in the release, I don't see how that's wasteful
godelski 5 hours ago [-]
For ages people have been saying "it's obvious" or "it's intuitive", without learning anything about design principles. Obvious to who?
People have been saying "we don't need docs", "nobody reads docs", "they'll figure it out". They experience the pain of a new employee onboarding but never treat this as a signal for the user experience. The user can't just walk over to the developer's desk. But luckily we got stack overflow, blogs, and Google. Somewhere someone explains it! The problem became finding it!
But hey, why do things the "hard" way when you can just throw money at the problem? Or better yet, dismiss the problem by calling the user an idiot.
You're right. What a time to be alive. We've created a world where this product is useful. And not even to because dark patterns exist, but simply because we've always wanted to avoid taking a step back and thinking or getting an outside view. Because it isn't obvious if a robot has to teach you what button to press. If it does, you probably should be ashamed (there are, as always, exceptions)
That said, I do think there's still a lot of utility you this. Those exceptions aren't uncommon. This'll help people learn programs like FreeCAD or Blender, or whatever. Where the complexity is naturally high. But also I do think we should recognize the silliness of many problems that this does solve that shouldn't be problems in the first place.
joquarky 2 hours ago [-]
> Obvious to who?
Garage sale signs are a good example of this. The person making the sign knows what it says, so they can "read" it from farther away than someone who doesn't already know what it says.
serf 6 hours ago [-]
i've read enough 'shlemiel the painter' ports across hundreds of systems and thousands of examples that I must quickly and coldly dismiss the precept that computers were ever used efficiently.
Anon1096 5 hours ago [-]
Now imagine what people back in the day thought about going from assembly to C. Or C to Java. Or, forgive me for even saying it, Java to Python.
sudo_cowsay 6 hours ago [-]
On the other hand, (while you made a completely valid point,) some people who are new to a field, like young aspiring computer scientists who don't know how to do some stuff and need a guiding hand, will find this tremendously useful. Also, technology and compute power has been growing a tremendous amount over the past 5-10 years (now we have 20-100 billion transistors in chips!). <-- enough to sustain this wild use of compute
Indeed, what a time to be able.
6 hours ago [-]
3 hours ago [-]
CamperBob2 3 hours ago [-]
If you look at the example animations, it's clear that Apple has brought this entirely upon themselves. Don't yell at some guys who are trying to fix it.
threethirtytwo 1 hours ago [-]
I’m really curious about the end game. When this technology completely saturates and AI becomes his mundane bullshit thing know one cares about but uses all the time. Like the next level above slop. I wonder what that world will look like.
sicktriple 13 minutes ago [-]
I also am curious of this, I only fear that it's going to turn into a concerted effort by those who have power to remove already existing capabilities from individuals and replace it with centralized infrastructure. Swapping out everyone's desktop for a think client hooked up to an agent and renting it back for a monthly fee for example, would be a strict downgrade when we already had the capability to do pretty much anything we could think of on the internet. Sure maybe it makes someone's workflow a b it more efficient, but I quite like having my own agency. I don't intend to hand it away some dork who owns a billion GPUs can rent it back to me.
6 hours ago [-]
redanddead 2 hours ago [-]
Human in the loop is the superior workflow!
/s
hn8726 11 hours ago [-]
I tried to read the "Does it need Screen Recording or Accessibility?" part, but it's slopped to the point I have no clue what it's trying to say. But if it can draw on top of permission prompts, what's stopping it from drawing box that hides the "decline" button and changing the "approve" button copy?
causal 10 hours ago [-]
I wonder if someday we will get to the point where Github repos are just markdown files describing the project and then you just let your own agent implement it because why the hell would I trust your agent's implementation?
jaggederest 6 hours ago [-]
Here's a couple examples, I've seen others as well:
I think it was more popular around the beginning of the year to mid-year
nater5000 9 hours ago [-]
Seriously. I imagine this will be the case soon enough in some shape or form.
Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.
Abimelex 8 hours ago [-]
A 100%! Just imagine you could describe every solution to a problem just using language! Of cause you would to make sure to avoid ANY missunderstanding, but therfore you could invent a language specified for avoiding ambiguity. I could imagine just to reduce the words to a very small corpus so everybody can remember it and have a very strict grammar so a program can effortlessly check correctness of its sentences.
drowsspa 15 minutes ago [-]
Well, I've always thought we are stuck in a local minimum when it comes to programming languages. Like, programming shouldn't be this hard and shouldn't require so much boilerplate after almost a century. If the actual important logic can be inferred pretty well from markdown specs, that means most of what we write in JavaScript, Go, Java is redundant.
Hopefully someone smarter than me is working with trying to develop a new programming language given what we know to be possible with LLMs
ccozan 7 hours ago [-]
I hope someone reads your post. I had to laugh hysterically when I reached the end. The level of sarcasm is too high.
joquarky 2 hours ago [-]
You have not experienced spec writing until you have written them in the original Klingon.
tempest_ 7 hours ago [-]
This only works while tokens are artificially cheap. The gravy train could come to an end at some point.
The open models that are chasing the frontier labs will stop being open once things slow down and there is less incentive to undercut the front runners. Time will tell if GPU compute gets cheap enough to run stuff locally.
zaik 7 hours ago [-]
This post makes me want to buy Nvidia stocks.
ccozan 7 hours ago [-]
/goal build this but better
theropost 8 hours ago [-]
Isn't it already kind of like that? Except the agent reads the code as if it's markdown. There's really no difference anymore, is there?
asdff 6 hours ago [-]
With enough model drift that won’t even work over time. These files would have to be pinned to the intended model version and that is either used directly or emulated with a faithful emulator in a larger model.
cootsnuck 5 hours ago [-]
I doubt that will be a problem. Models are different enough right now. If you can write a spec with enough detail and constraints today to get valid and comparable output from say GPT-6 Astra, Opus 5.5, DSV4F, Kimi K3, GLM 5.3, etc... Then I think there's a good chance that whatever SoTA coding LLMs everyone is using 3 years from now will also be able to implement that same spec.
Again, I think it heavily depends on if people are writing comprehensive specs with sufficient detail.
asdff 3 hours ago [-]
That depends on if you are really getting what you expect from the spec or you are actually relying on undefined behavior for your expected output from the spec. This is why we often pin software library versions: we may very well be relying on unintended or a bugged behavior to get out expected output.
rrr_oh_man 6 hours ago [-]
llmpm
glitchc 7 hours ago [-]
Sounds like Github is optional at that point. Why not just use Medium or WordPress?
asdff 6 hours ago [-]
Just tell claude to share it with other claude users
lstodd 6 hours ago [-]
reddit! use reddit!
m-s-y 7 hours ago [-]
ding ding! we have a winner!
fennecfoxy 8 hours ago [-]
Eh I think it'll be more like Gibson's defensive ICE in that ICE is "my swarm of agents scans your code for nasties, then compiles it from source".
glitchc 7 hours ago [-]
Neither party owns their agents though. Is it simply then a matter of "my subscription is better than your subscription"?
pydry 8 hours ago [-]
First you'd have to stop the agents from flagging hundreds of irrelevant nasties.
I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.
cyanydeez 10 hours ago [-]
if all your models are in the cloud, why would you trust anything your agent builds
jerf 10 hours ago [-]
That sounds cynical today, but that's just because the AI models are currently outrunning enshittification. I'm already pondering personal plans about what to do when that turns around. We haven't seen enshittification yet that is going to be like the enshittification of AI. It may even deserve a new term of its very own, it's going to be such a big problem. The AI companies are leaving a lot of value on the table to entice us on to their systems but at some point that's going to turn around.
cyanydeez 9 hours ago [-]
The existential danger exists today: you're providing your entire business toolchain to models in the cloud, owned by businesses who are for profit entity. Even if the model itself doesn't care about your data, the business model does.
Everything your building is now streaming through a third party who could steal it, the same way Amazon steals open source software and repackages it.
All you need to do is look at palentir's forward deployed engineers. Once you let a third party observe your processed, you're cloneable and replaceable _as an entity_.
I think we can look at our own human psychology as it developed into social groups. If our brains were obviously deterministic, it would be easy for people to prey upon us via structured inputs. Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.
So these AI companies are definitely more dangerous than just the "AI is super intelligence". They're basically a intelligence platform everyone is wiring into.
That's today. Tomorrow, they'll be run by an MBA which is the enshittification.
mat_b 4 hours ago [-]
> Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.
Not really. Ants are superorganisms, they share more than half of their DNA with each other. You could think of each ant as being like a cell in a human.
jack_pp 9 hours ago [-]
software businesses are not like watches (recently saw a youtube video about fakes being made virtually identical at 10% of the price).
If you can make a facebook clone it doesn't mean you can steal facebook's lunch. there are exceptions of course.
hannasanarion 9 hours ago [-]
The fact that the utility doesn't give it the ability to do that?
It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.
So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?
For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.
SwtCyber 10 hours ago [-]
Nothing stops it, except the fact that the agent is already executing arditrary code in your shell. If it's malicious, it'll just steal your shh keys directly instead of bothering with button masking
smugglerFlynn 8 hours ago [-]
Don't tell me you are reading readmes with your eyes in 2026. Blasphemy!
tkdb 10 hours ago [-]
...and there we have it. Slop assumes a verb form.
You know the rules. Now that it's AI slopping it's new and innovative.
tottenhm 8 hours ago [-]
Some words reverberate.
trollbridge 9 hours ago [-]
Indeed, I’ve heard it on the context of feeding animals my entire life.
nedt 18 minutes ago [-]
Funniest thing about this is how the arrows look like. They aren't just straight or one clear curve. If you have ever seen what Franz draws in workshops, this mix of art and a children drawing, that's how the arrows look like.
I'm wondering if the AI has been trained on Franz. Which would also explain why the Terminator ends up with Manner Schnitten.
internet101010 3 hours ago [-]
The worst trend in UX in the last decade is the endless "Got it!" popups and feature notifications that distract the user from what they were trying to do.
I don't know why anyone would ever willingly want this.
tangotaylor 9 hours ago [-]
"It is an arrow, so we spent an unreasonable amount of time on how it looks."
Brilliant. This is exactly the kind of content I seek when I visit Hacker News.
Truly art.
socializer 8 hours ago [-]
Have you actually looked at this? In the first screenshot, three of the five arrows are obviously and badly misaligned, have incorrect labels, and it's just a big tangled mess (edit: the author stealth-replaced the screenshot, but the original is here: https://raw.githubusercontent.com/franzenzenhofer/big-arrow-...). The LLM-generated text may be saying one thing, but there's clearly zero effort spent on... anything. There's no punchline, there's no aesthetic angle, there's no conceivable purpose.
It's just the epitome of living in an era where code is so cheap to generate that we don't invest any effort into even fully fleshing out the idea. It's just "Claude, make me a program that draws arrows on screen". And HN, I guess, is dominated by two groups: one still stuck in the mentality of "code = effort" and trying to recognize that; and another stuck in the mentality of "AI = cool".
threethirtytwo 59 minutes ago [-]
It’s inevitable. The future is “prompt = effort”
Those who code will be left behind. I code. I don’t like it. But I have to face reality.
kogus 6 hours ago [-]
My first reaction was similar to the reaction I'd have if you told me that cockroaches had learned to unlock doors and stand on their hind legs. But then I thought about accessibility, and the ways this could be used to help technologically illiterate or disabled people, and I thought again.
Do you remember when PCs used to come with a completely soup-to-nuts tutorial that would talk to you like you had never seen a PC before? Things like this: https://www.youtube.com/watch?v=3ScS4OYDfHE
This kind of baked-in interactivity could really help in a training or disability context.
usrbinbash 10 hours ago [-]
SO the point of this is ... what exactly?
A big arrow to an interface element which ... has a label that explains what it does?
So...a label for a label?
hannasanarion 9 hours ago [-]
It seems silly but I think there's a good use case for this:
Helping people deal with bad UX.
Everybody who has worked in a traditional corporate environment has encountered poorly tutorialized, poorly labeled, counterintuitive user interfaces where the only way to know how to get something done is to have somebody sit over your shoulder and tell you which buttons to click.
Oh, you want to figure out which OIM entitlements your coworker has that you don't to debug an access issue? You just need to know that you have to click the "compliance" button to get there. Why is it called "compliance" when what it actually does is let you export comparisons from the user tables? Because fuck you that's why. Memorize it.
An AI assistant that can read the procedures for you, or sometimes even read the source code, or the website code, can show you how to accomplish something that the UI on its own fails to do.
wartywhoa23 5 hours ago [-]
Wait, I was under impression that UX was to be sorted out by AI just like the code is?
voidUpdate 10 hours ago [-]
It's so your agent can make a big arrow on the screen saying "click this, human" so that it can keep doing things
andyfilms1 10 hours ago [-]
I love living in the future!
mistersquid 9 hours ago [-]
I ran two experiments which require the `claude` CLI tool be installed.
For the first experiment, Claude stepped me through how to edit its settings.json to remove warnings from its configuration: a little clunky as I had to respond to Claude's CLI prompts to get it to auto run.
In the second experiment, Claude walked me through how to use Apple's Compressor.app to speed up a video. It walked me through dragging the file into compressor and clicking on the appropriate UI elements to do this.
So, one use case for the bigarrow skill enables Claude to step users through a relatively complex UI to accomplish a task.
There are likely more efficient ways to accomplish the tasks at hand (e.g. ffmpeg for the second experiment), but this use case is interesting if using a GUI tool is required.
cobbal 7 hours ago [-]
Finally we've built the reverse centaur from Cory Doctorow's classic novel "Don't Become the Reverse Centaur"
ghm2180 10 hours ago [-]
Glad you asked. Pointing my aging aunt to the right place on the screen to click without having to take control of her laptop, of course.
arshxyz 12 hours ago [-]
The README is geared towards technical people (complete with the HN screenshot) but when I see a tool like this all I can think of is how helpful this would be for my mom when I'm trying to tell her how to download and print a document over the phone
conception 11 hours ago [-]
Honestly the giant arrow annotation Zoom has makes it worth any amount of money compared to the competition.
prmoustache 5 hours ago [-]
Except your mom just have to tell her AI agent to download and print that document.
She doesn't have to call you anymore.
isoprophlex 12 hours ago [-]
Literally unusable as it is. Some minimal extra features this would need:
- rainbow dripping arrows
- angrily pointing arrows
- flame-surrounded text boxes with particle effects
- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings
This is where the agent says did not understand sarcasm coded and shipped features
DonHopkins 11 hours ago [-]
Last time I accidentally said something sarcasticly over-ambitious to an LLM, it shipped this popup callout tooltip feature on a PDP-7 Type 340 vector graphics display emulator that shows you the meaning of the drawing you're pointing at, as well as the address of the instruction that drew it.
At the risk of stating the obvious - let's not do that, the goal of this repo is to be useful and not to give agents the power of the `<blink>` tag.
voidUpdate 10 hours ago [-]
> ""I need you, and you're making coffee." --say reads the sign aloud. Your Mac will literally call you back to your desk."
It's already got the power to be obnoxious at you
koalacola 11 hours ago [-]
Oh dear, they were making a joke.
yen223 11 hours ago [-]
if only there was a way to make a subtle thing obvious
isoprophlex 11 hours ago [-]
such as... angry flaming rainbow textboxes and arrows?
ale42 11 hours ago [-]
I thought that the dripping rainbow ones were enough. Maybe you have to ask for rainbow unicorns flying on the screen.
jaapz 10 hours ago [-]
the repo is one big joke, of course they should add this
thih9 8 hours ago [-]
[dead]
DonHopkins 11 hours ago [-]
For the humor impared, it would also be useful to have a colorful animated "WHOOSH" overlay with sound effects for every time a deadpan joke goes over your head. ;)
Maybe isoprophlex will add that to his PR!
ipsod 12 hours ago [-]
under_construction.gif
isoprophlex 11 hours ago [-]
just submitted the airhorn PR; ~second rainbow arrow slop grenade incoming~ BOOM slop cannon fired
I would pay good money for something like this on iPad. It wouldn't even need to be agent-driven, just a big "I want to make a bank transfer" button, that would launch the correct app and guide through the interface.
That would be a godsent for those of us with aging parents.
ghm2180 11 hours ago [-]
> That would be a godsent for those of us with aging parents.
Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.
It's Time to redesign the apple genius bar for the modern aging boomer using this.
I'm just as impressed by this as any other LLM-generated project.
priyashunt 7 hours ago [-]
Very VERY USABLE FOR old PEOPLE. I would pay good money for something like this on iPad!
vessenes 12 hours ago [-]
Interesting. When I read the headline I imagined this would be a sort of thinking trace booster -- letting the agent focus its own attention on different parts of the screen. But this is cool in a different way. I bet agentic harnesses would find it useful for communicating with other agents / themselves as well.
alansaber 12 hours ago [-]
This might be goofy, but it underscores that there's potential for more visual agent UIUX than reading off a sidebar/opening modals.
melvinroest 11 hours ago [-]
[dead]
melvinroest 11 hours ago [-]
My message to the world is that LLMs should be able to point anything they see in the application they're in or even the whole computer (if you give it that kind of access).
For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.
We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?
FinnLobsien 12 hours ago [-]
This could be great for documentation. Screenshots in docs are frequently useless because they show me a screen and say "click X" where I still have to search X visually. And I could just to dhat in the other tab I have open.
peaxkl 11 hours ago [-]
It doesn’t make sense that you have to read a whole article and then still search for the buttons in the UI afterwards. And with longer articles, you always have to keep the article open next to your product to follow the whole flow.
What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?
If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.
inanutshellus 12 hours ago [-]
The first example (of HN) is the one that feels like it has the most potential to me.
"Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".
Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?
ravila4 10 hours ago [-]
I think this would’ve been very handy to me a couple years ago when I was learning to use Blender and asking LLMs for help performing certain actions like “how do I display the normals of all the vertices in my mesh?” I spent a lot of time trying to figure out which button the model was talking about.
alexpotato 9 hours ago [-]
> Arrows have existed since roughly the Paleolithic.
Whenever I use web tools that don't have "deep linking" [0], I love to throw out this quote:
"Have we thought about using hyperlinks? They are these SUPER useful things that were invented in Switzerland back in the 1990s."
I want this for the visualization of very difficult to solve application evolution.
This is for problems which have no known solution and are too complicated for an agent to figure out on its own.
My current problem: can alphafold and related tools figure out the function of a gene that so far is unknown.
Think of the monte carlo tree search in the alphago explanation videos.
What if there was visualization that showed how the policy and value models worked so you can spot their flaws.
I want the agent to build an attempt at a solution,
then build a visualization of it,
then run the solution and show me what is it up to.
I want to it pause and explain its state so I can see exactly where and why the solution fails.
emigre 1 hours ago [-]
This reminds me of good old Clippy.
dr_kiszonka 8 hours ago [-]
I take this opportunity to shame GitHub for completely ignoring the mobile experience in their own Android app. The project's README is pages upon pages of largely blank space. In general, the app does not render mermaid diagrams and does not allow for zooming in, so smaller pictures are unusable. Even if you access pictures directly in a repo's source, you still can't zoom in. iPython Notebooks, which are extremely common in data science, are an "unsupported file type" and are not rendered.
(OP, nice project! Sorry for my rant.)
bel8 7 hours ago [-]
I'm forced to use GH mobile app because of 2FA.
But it's a subpar experience compared to just opening github on the browser.
githubnuoo 7 hours ago [-]
it was driving me mad too. you can disable github.com opening in the github app: app-info/set-as-default/open-supported-links
m-s-y 11 hours ago [-]
Genuine question…
Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.
IanCal 10 hours ago [-]
There is a skill, but you don't have to use it, though you'd have to explain how to use the app.
> burning tens of thousands of tokens in context.
~1400?
> Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”
Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.
raviinits 7 hours ago [-]
[flagged]
ballofrubber1 10 hours ago [-]
With a skill claude will know when to use it without you specifically prompting it.
TekMol 12 hours ago [-]
Swift, Shell, Python and Objective-C
Does one need 4 programming languages to draw something on a mac?
sitzkrieg 12 hours ago [-]
welcome to zombocom. err i mean modern HN :-(
ElijahLynn 4 hours ago [-]
This is going to be really useful actually, nice work! Can't wait for it to be on Linux too!
ghm2180 11 hours ago [-]
Man, Ive lost count of How many times have I had to repeat this over the phone to my parents; The repo has all the punch lines in the README
> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".
> Remote help. "No, the other gear icon." Point at it instead of describing it.
xyzsparetimexyz 12 hours ago [-]
Seems like a pretyu useful way to help infants use desktop computers
ipsod 11 hours ago [-]
Have you met users?
iandanforth 7 hours ago [-]
Pointing is surprisingly overlooked in a lot of tools and I look forward to using this. It reminds me of what I thought was a killer feature on AnyBots telepresense robots, they had a laser you could use to point at things while piloting the robot.
LoneRanger1024 8 hours ago [-]
I think this makes a lot of sense for tutorials, but permission prompts need different treatment. Besides pointing to a button, the assistant should explain what clicking it will do, so the user can make their own decision.
tilemarch 10 hours ago [-]
An arrow on the screen solves “where do I click”; it doesn’t solve “should this happen”
To me this feels like it takes away from what the human is supposed to do (read, understand the consequences of the action, then.. consent or abort)
There is a reason your AI Agent won't automate these clicks for you
ceroxylon 6 hours ago [-]
This is like the alerts in new cars that remind you to check for other passengers when you exit... helpful, but a worrying sign of the times.
user- 6 hours ago [-]
Seems great for scammers targeting old people tbh
flr03 11 hours ago [-]
That would have been handy 20 years ago to point to that one valid 'Download' button.
Maken 11 hours ago [-]
How could you tell apart the legit arrow pointing to the download button from all the fake ones?
pimlottc 10 hours ago [-]
Even the example arrows in the first screenshot are wonky and unnaturally weird.
code_duck 6 hours ago [-]
This is definitely not something I want, ever. Maybe someone could be helped by it I guess.
antonyragleap 8 hours ago [-]
Clever approach for agent debugging. Visual cues > logs when multiple agents run. How do you handle overlapping annotations?
ForHackernews 12 hours ago [-]
"and they keep hitting the same wall, the part that only a human may do"
Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.
franze 12 hours ago [-]
Claude refuses to do certain actions (enter passwords, change security settings, create new accounts on external services) even in Yolo mode running as sudo. (I tested it all on its own mac machine)
gwerbin 12 hours ago [-]
You can add custom auto-mode classifier rules and even disable the built-in ones, if you want to live on the edge like this.
jaapz 10 hours ago [-]
isn't `--dangerously-skip-permissions` enough?
franze 6 hours ago [-]
no, it is not the permissions that are holding the model back from above cases, it is the system prompt and / or the model itself.
sitkack 8 hours ago [-]
--norms-taboos-off
12 hours ago [-]
ex-aws-dude 11 hours ago [-]
If you can’t even take the time to understand what you’re clicking why even go through the formality of “approving”
harrouet 10 hours ago [-]
There is no limit to burning tokens :)
priyashunt 7 hours ago [-]
THERE IS
amelius 10 hours ago [-]
Because AI can paint pelicans on bicycles quite well, and not arrows?
lapestenoire 12 hours ago [-]
I love it.
Retr0id 11 hours ago [-]
Reminds me of something from Idiocracy (2006)
okasaki 9 hours ago [-]
Look at what Windows and Mac users need to mimic a fraction of our power
wartywhoa23 5 hours ago [-]
Next up: Spare yourself from formulating any prompts, let our SI agent control our thin client by asking you leading questions that you answer by drooling for yes and picking the nose for no!
erickhill 8 hours ago [-]
I don't know why exactly but something about the arrows feel repulsive and oddly gross.
Reminds me of Idiocracy. Are we idiots yet? This is so that AI agents can "guide" a human, directing the meat bot more easily. Telling the meat bot where to click and what to do. The entire "what to think" step isn't even necessary when there is no thought. Because thoughts and decisions are annoying. It's astonishing how much people are already "algorithmically" steered so they don't have to turn on their rational thinking.
seems useful for most of the folks that just want to do stuff without going through (over)complicated UIs...
But I feel dumber just by looking at it. If this is how this will look like in 10 years, why to make desktop at all. Just connect mic and speaker to your PC or talk to the phone: 'I need new pair of socks. Order 10 for me. In your favourige color it will be 21.37. Should I charge your credit card?'.
It is not like most of the people enjoy computers. I am pretty sure they do not. They just need them to operate systems they need: government websites, banks, maps, restaurant menus etc. If some agent will do that for them, why bother looking at screen at all? Rich people have it with their own personal assistants.
sehw 10 hours ago [-]
Fuck you, I won't do what you tell me.
mococa 11 hours ago [-]
Another slop project on front page.
Gareth321 10 hours ago [-]
While it's vibe-coded, I'm immensely enjoying the creativity which AI has enabled. These apps would never have been built before, and I find it so fun to see how AI is being used to provide (semi) useful things for people who need it.
I found this awesome vibe-coded Pixel News Network project recently. How fucking cool is this?? Would never have been created otherwise. https://pnn.watch/
8 hours ago [-]
mbty 7 hours ago [-]
Meh, would have been cool-ish a few years back, now everyone and their grandmas could build something similar given enough tokens. Hard to be excited when similar gadget projects are a dime a dozen. I also don't think that AI use enables much in the way of creativity. It allows people to delegate their judgment to the model, which defaults to humdrum middle-of-the-road choices. Only the very high-level decisions remain personal, leading to a lack of texture past the first glance, so to speak. Not saying that it cannot be used with intentionality, just that it rarely is.
I actually somewhat enjoyed the times when vibed projects were a broken mess, lots of accidental comedy in that era.
quikroofficial 11 hours ago [-]
best
coldbootHq 9 hours ago [-]
[flagged]
SwtCyber 10 hours ago [-]
[flagged]
colinmarc 12 hours ago [-]
[flagged]
mouldloft 9 hours ago [-]
[flagged]
ProofHouse 6 hours ago [-]
[flagged]
einpoklum 12 hours ago [-]
More LLM-authored items about LLM slop.
inanutshellus 12 hours ago [-]
As long as it has value... I'll allow it.
~guywithnopowertodisallowit
DonHopkins 11 hours ago [-]
[flagged]
nixosbestos 12 hours ago [-]
What a time to be a radical centrist - the AI haters seem out of touch, the AI thought leaders can't stop huffing their farts and being condescending, and somehow this is on the top of HN. What a silly time.
VCFundedGenYer 11 hours ago [-]
Please don't submit low quality content like this to HN.
When we got a form of such search, it searched all of the internet alongside with the functions of your software on your computer. You had to tell it manually what software you have and when it found something you had to fish out the description of what to click to get the functionality you are seeking and then click through some dumb ui to actually invoke it.
Agents are the first thing that can lead you directly from "I know what I want my computer to do." to "Actually doing it."
Success of agents is founded on the profound and sustained failure of all of the software designers and developers ever.
apropos ?
Obviously I'd prefer to ask specific questions to my expert human friend sitting at the next desk over, but sadly they don't exist.
I just hope there is a future where this tech exists without such big downsides.
Even the data centre using this probably uses less net energy, and costs less, to build one of these dumb arrows than the aggregate sum of the energy used to keep you alive while you Googled for the answer. (Including heating/cooling the building, powering the equipment growing the food you eat, etc.)
Nonetheless why not let an LLM make a guide and use that in the release, I don't see how that's wasteful
People have been saying "we don't need docs", "nobody reads docs", "they'll figure it out". They experience the pain of a new employee onboarding but never treat this as a signal for the user experience. The user can't just walk over to the developer's desk. But luckily we got stack overflow, blogs, and Google. Somewhere someone explains it! The problem became finding it!
But hey, why do things the "hard" way when you can just throw money at the problem? Or better yet, dismiss the problem by calling the user an idiot.
You're right. What a time to be alive. We've created a world where this product is useful. And not even to because dark patterns exist, but simply because we've always wanted to avoid taking a step back and thinking or getting an outside view. Because it isn't obvious if a robot has to teach you what button to press. If it does, you probably should be ashamed (there are, as always, exceptions)
That said, I do think there's still a lot of utility you this. Those exceptions aren't uncommon. This'll help people learn programs like FreeCAD or Blender, or whatever. Where the complexity is naturally high. But also I do think we should recognize the silliness of many problems that this does solve that shouldn't be problems in the first place.
Garage sale signs are a good example of this. The person making the sign knows what it says, so they can "read" it from farther away than someone who doesn't already know what it says.
Indeed, what a time to be able.
/s
https://github.com/seb3773/ntfs-repair-rfc
https://github.com/Kotivskyi/screenshot-tool
I think it was more popular around the beginning of the year to mid-year
Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.
Hopefully someone smarter than me is working with trying to develop a new programming language given what we know to be possible with LLMs
The open models that are chasing the frontier labs will stop being open once things slow down and there is less incentive to undercut the front runners. Time will tell if GPU compute gets cheap enough to run stuff locally.
Again, I think it heavily depends on if people are writing comprehensive specs with sufficient detail.
I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.
Everything your building is now streaming through a third party who could steal it, the same way Amazon steals open source software and repackages it.
All you need to do is look at palentir's forward deployed engineers. Once you let a third party observe your processed, you're cloneable and replaceable _as an entity_.
I think we can look at our own human psychology as it developed into social groups. If our brains were obviously deterministic, it would be easy for people to prey upon us via structured inputs. Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.
So these AI companies are definitely more dangerous than just the "AI is super intelligence". They're basically a intelligence platform everyone is wiring into.
That's today. Tomorrow, they'll be run by an MBA which is the enshittification.
Not really. Ants are superorganisms, they share more than half of their DNA with each other. You could think of each ant as being like a cell in a human.
If you can make a facebook clone it doesn't mean you can steal facebook's lunch. there are exceptions of course.
It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.
So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?
For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.
I'm wondering if the AI has been trained on Franz. Which would also explain why the Terminator ends up with Manner Schnitten.
I don't know why anyone would ever willingly want this.
Brilliant. This is exactly the kind of content I seek when I visit Hacker News.
Truly art.
It's just the epitome of living in an era where code is so cheap to generate that we don't invest any effort into even fully fleshing out the idea. It's just "Claude, make me a program that draws arrows on screen". And HN, I guess, is dominated by two groups: one still stuck in the mentality of "code = effort" and trying to recognize that; and another stuck in the mentality of "AI = cool".
Those who code will be left behind. I code. I don’t like it. But I have to face reality.
Do you remember when PCs used to come with a completely soup-to-nuts tutorial that would talk to you like you had never seen a PC before? Things like this: https://www.youtube.com/watch?v=3ScS4OYDfHE
This kind of baked-in interactivity could really help in a training or disability context.
A big arrow to an interface element which ... has a label that explains what it does?
So...a label for a label?
Helping people deal with bad UX.
Everybody who has worked in a traditional corporate environment has encountered poorly tutorialized, poorly labeled, counterintuitive user interfaces where the only way to know how to get something done is to have somebody sit over your shoulder and tell you which buttons to click.
Oh, you want to figure out which OIM entitlements your coworker has that you don't to debug an access issue? You just need to know that you have to click the "compliance" button to get there. Why is it called "compliance" when what it actually does is let you export comparisons from the user tables? Because fuck you that's why. Memorize it.
An AI assistant that can read the procedures for you, or sometimes even read the source code, or the website code, can show you how to accomplish something that the UI on its own fails to do.
For the first experiment, Claude stepped me through how to edit its settings.json to remove warnings from its configuration: a little clunky as I had to respond to Claude's CLI prompts to get it to auto run.
In the second experiment, Claude walked me through how to use Apple's Compressor.app to speed up a video. It walked me through dragging the file into compressor and clicking on the appropriate UI elements to do this.
So, one use case for the bigarrow skill enables Claude to step users through a relatively complex UI to accomplish a task.
There are likely more efficient ways to accomplish the tasks at hand (e.g. ffmpeg for the second experiment), but this use case is interesting if using a GUI tool is required.
She doesn't have to call you anymore.
- rainbow dripping arrows
- angrily pointing arrows
- flame-surrounded text boxes with particle effects
- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings
EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...
https://hyperties.org/cabinet/symelec/
PIXIE and FORTH on PDP-7 Cabinet Emulator with Type 340 Vector Graphics Display:
https://www.youtube.com/watch?v=lo8kdY-5i6c
It's already got the power to be obnoxious at you
Maybe isoprophlex will add that to his PR!
That would be a godsent for those of us with aging parents.
Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.
It's Time to redesign the apple genius bar for the modern aging boomer using this.
For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.
We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?
We built something to help with that [1].
[1] https://www.happysupport.ai/en/in-app-messaging
What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?
If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.
"Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".
Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?
Whenever I use web tools that don't have "deep linking" [0], I love to throw out this quote:
"Have we thought about using hyperlinks? They are these SUPER useful things that were invented in Switzerland back in the 1990s."
0 - https://en.wikipedia.org/wiki/Deep_linking
(OP, nice project! Sorry for my rant.)
But it's a subpar experience compared to just opening github on the browser.
Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.
> burning tens of thousands of tokens in context.
~1400?
> Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”
Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.
Does one need 4 programming languages to draw something on a mac?
> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".
> Remote help. "No, the other gear icon." Point at it instead of describing it.
To me this feels like it takes away from what the human is supposed to do (read, understand the consequences of the action, then.. consent or abort)
There is a reason your AI Agent won't automate these clicks for you
Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.
https://laughingsquid.com/yahoo-neon-billboard-in-san-franci...
I remember taking your SEO course many moons ago.
I absolutely loathe this phrasing now. I don’t even know what part of it is good or notable.
PSIBER Space Deck and Pseudo Scientific Visualizer Demo:
https://youtu.be/_fqCeuue5Ac?t=213
The Shape of PSIBER Space: PostScript Interactive Bug Eradication Routines — October 1989:
https://medium.com/@donhopkins/the-shape-of-psiber-space-oct...
But I feel dumber just by looking at it. If this is how this will look like in 10 years, why to make desktop at all. Just connect mic and speaker to your PC or talk to the phone: 'I need new pair of socks. Order 10 for me. In your favourige color it will be 21.37. Should I charge your credit card?'.
It is not like most of the people enjoy computers. I am pretty sure they do not. They just need them to operate systems they need: government websites, banks, maps, restaurant menus etc. If some agent will do that for them, why bother looking at screen at all? Rich people have it with their own personal assistants.
I found this awesome vibe-coded Pixel News Network project recently. How fucking cool is this?? Would never have been created otherwise. https://pnn.watch/
I actually somewhat enjoyed the times when vibed projects were a broken mess, lots of accidental comedy in that era.
~guywithnopowertodisallowit