Docker Agent
github.com/docker
[5 comments hidden]
How is "no code required" a selling point when we all have agents that can almost instantly puke out all the mostly-correct code you could want?
[hidden]
[hidden]
[29 comments hidden]
All the cool kids have one!
[11 comments hidden]
I say this as a gopher who also has a custom harness written in Go, it's a real challenge compared to the TS setups, but TS plugins are also a security concern
[2 comments hidden]
I haven't given this much thought (I also just started with golang last year) but I'm assuming the only real path here is to provide extension SDKs in Lua or other languages whose interpreters have been implemented in go.
I think that's what grafana's k6 [1] does with a JS interpreter
[hidden]
Go has the plugin package, but I haven't actually seen anyone use it (outside of toys and blog posts)
https://oneuptime.com/blog/post/2026-01-25-plugin-system-go-...
[2 comments hidden]
[4 comments hidden]
[3 comments hidden]
It may work for some things, but you really want to be wired into that system to do the more interesting things.
[hidden]
jk its trash
[11 comments hidden]
Vue and svelte are easy to use and learn/adopt. But look what won. React.js.
So whatever the ai agent with the vast user basae will win regardless.
---
I know it's a slippery slope, but given you brought up JS framwork analogy, I had to fall into it
[2 comments hidden]
Now Angular v1, man… if I ever see another digest loop error in my life, I will scream
[hidden]
[5 comments hidden]
Cursor is easily the worst experience I've had with an AI harness, yet it was acquired for billions of dollars in spite of being a middleman tied to a ripoff of VS Code. None of the colleagues of mine who were touting it last year are still using it. But if you can make yourself look like the next big thing, you'll get money thrown at you. Just look at Omarchy.
A good experience necessarily takes time and careful thought. Nobody these days can do that while also remaining relevant, or even surviving.
[hidden]
I kinda like cursor, mostly because it provides a nice review UI and I can use whatever model I want.
Getting away from all the Claude-speak has been a revelation, and fortunately Anthropic helped me here by hyping up GLM 5.3. Like, if there's an open weight Fable competitor, then why wouldn't I use it?
[hidden]
"Uh, OK, so how do I get this thing working? Where are the docs?"
"Well, first you install its unique framework..."
[7 comments hidden]
If, like me, you couldn't find any security-related info on the linked page.
[hidden]
[5 comments hidden]
If you don't want to use their harness then you wouldn't use docker agent and instead use their `sbx` cli to run the harness of your choice (claude, codex, pi, etc).
I think it will be great if Docker can get people used to using secure VMs. I am developing a similar project (still a work in progress): https://github.com/gregwebs/agent-vm
[3 comments hidden]
[hidden]
[6 comments hidden]
[5 comments hidden]
[hidden]
[15 comments hidden]
[6 comments hidden]
Having it docker branded, I could understand. It’s confusing but the docker brand is strong.
But exposing it as a docker subcommand is very confusing to me.
[hidden]
But it evolved differently and although it runs very well in docker containers and docker sandboxes, it also runs very well anywhere else.
[hidden]
It is called 'Docker' as the company, and it has nothing to do with the container technology.
[hidden]
Great tech. Not a business.
[3 comments hidden]
Does anyone remember all these web portals from 2000? What people mostly wanted was a search engine but they added everything else and lost focus.
[2 comments hidden]
But currently uv is good enough to create isolated, repeatable environments. And it isn't limited to Linux like standard Docker. So when I want to test small samples on my local PC's GPU (Windows or mac) before running on beefier hardware, I can do that easily.
[hidden]
[10 comments hidden]
[6 comments hidden]
[5 comments hidden]
I tried not to do anything too crazy, but even use using an API key in the docker agent natively was getting me sessions that would hang and couldn't be recovered.
[3 comments hidden]
[2 comments hidden]
[hidden]
I work at another big tech company, and up until recently that would’ve been true for us as well (but it’s changed now, we partnered with OpenAI)
[hidden]
[hidden]
[3 comments hidden]
[hidden]
[5 comments hidden]
Yep, open ai sol model wrote that.
[3 comments hidden]
[hidden]
There was a brief point in time at the beginning of the internet where if you saw an image online you could be kind of confident that it hadn't been manipulated significantly... But somewhere around the mid to late 2000s it became prudent to just assume any image you've seen online got at least a blemish filter pass through something like Photoshop.
Same thing now with written text. Sometimes you'll be a little bit more or less sure than an llm produced the output but never able to say 100% confidently.
[hidden]
[hidden]
looks like a weird abstractions. why do i need this if i have modal.com, e2b, cloudflare that use original docker + some toolings around + way to run + isolations on network level.
[hidden]
[2 comments hidden]
Something like a containerized sudo ACL for commands and for network credential , scoped by oauth scope
That way you could truly establish boundaries for an agent before it’s launched , and it could request additional permissions during execution .
[hidden]
[2 comments hidden]
[hidden]
At the end, the link with docker compose is just the yaml form. Which by the way can be replaced with hcl files or Go code.
[hidden]
Have you considered making a cloudy, managed version?
[hidden]
[hidden]
[2 comments hidden]
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
just use and customise Pi.
Nothing else even comes close.
[hidden]
[hidden]
I'm not sure why anyone would want this. Is there something this actually does better than any other harness? Do current harnesses suck that bad at orchestrating things like Docker? If I do all my work inside a Linux VM and tell an agent inside of it what I need done, it figures out everything, including container orchestration.
The very job that AI gents are meant to do should mean that most of the documentation for this Docker Agent is obsolete/unnecessary.
[3 comments hidden]
Nowadays there are only two justifiable public languages you should be using for anything a) C++ b) Rust
Rust itself is dubious due to abysmal compile times and the fact that the only guarantees it gives you are of security (aka a skill issue). Modern LLMs write C++ that is as safe if not safer than Rust. Any other language choice is objectively wrong.
* For those inquisitive enough the keyword "public" is doing all the heavy lifting here. Pretty much everyone should be developing and using an in-house DSL at this point.
Olscore[38 comments hidden]
I've been working on Pullboard and just open sourced it: https://github.com/pullboard-dev/pullboard
To me, orchestration is not really the core problem. The bigger challenge is handling agent coherency over long periods of time, (drift) so my solution is a living forum-like system where agents shout to one another, document work in items, and follow doctrine handed down by the developer and work against a known specification that's checkable. I have used this workflow in several large projects that MUST be correct, and it works for me. Currently cleaning it up for other users.
pitched[30 comments hidden]
jbp[18 comments hidden]
redhale[13 comments hidden]
The vanilla out of the box product is actually quite close to old Trello, in simplicity. But there are so many ways to fuck it up and overcomplicate it, and so many companies proceed to fuck it up and overcomplicate it.
greggh[hidden]
saguntum[2 comments hidden]
I have learned recently that I don't inherently hate JIRA, I hate our JIRA and byzantine process flows.
I think because other task trackers like Trello are simpler and have less config options (per my last use, long ago, pre-Atlassian-acquisition), it's impossible knock on wood to get into the antipatterns I've seen in those tools.
Ironically I think one of the main goals of management I have seen consistently over time - ensuring tickets are accurate, up to date, and reflect reality - would end up much more true if the task tracking tools were just much simpler, with fewer fancy features.
ionetan[hidden]
greysonp[3 comments hidden]
tomashubelbauer[hidden]
NBJack[hidden]
makeitdouble[hidden]
So what they really hate is their company. Which is usually fair and understandable.
zerkten[hidden]
I think what you describe is a problem for a lot of people, but what isn't acknowledged is that some people don't want any kind of process, or rather, they don't want to deal with any kind of friction. These people don't like Jira without the customizations. They don't like equivalent tools.
Given these tools regiment the world to some extent, many programmers are fine with the minimal friction. However, you have big problems when people who don't value any process get into positions of authority. This is also often a problem when you interface with other units of an organization who don't understand the value of a process.
devmor[4 comments hidden]
It’s business software for business people that need to be guided through SOP because they have so many different workflows, not engineers that have streamlined work processes to reduce mental load on non-engineering work.
ionetan[3 comments hidden]
devmor[2 comments hidden]
Unfortunately I work for a very large corporation and don’t have the pull to get something like Epiq brought on - but I am definitely going to evaluate it for my non-FTE work.
As a side note - it was impossible to find your project through search engines. I only found it by searching on github itself.
ionetan[hidden]
Zambyte[3 comments hidden]
[0] https://github.com/ankitpokhrel/jira-cli
conception[2 comments hidden]
alemanek[hidden]
x-complexity[hidden]
Need a specific feature? Plug that in later when you actually need it.
pea[hidden]
Chat is great for novel open-ended work, but is terrible at concurrency - i.e. managing > 1 agent runs. As agents get cheaper, I suspect we'll see more more bounded processes where you want to run lots of items through them at once, and thus Jira-like things... (we ended up having this problem and building something like this, albeit more medieval themed [0])
We also found the other piece that Jira-style solves (and that chat makes super hard) is multi-player. I've yet to see a good implementation of a multi-player chat solution to agents working concurrently; and we found that an async model is the only sane solution to this.
[0] https://pumpup.com
jredwards[hidden]
https://yegge.ai/gastown
blitzar[hidden]
xgb84j[4 comments hidden]
dummydummy1234[hidden]
wccrawford[hidden]
ngruhn[hidden]
dionian[hidden]
toddmorey[2 comments hidden]
pitched[hidden]
ogou[hidden]
j_bum[2 comments hidden]
> gate: failing
Gave me a good chuckle. Cool idea though, and I like the records. Reminds me a bit of the OpenAI hacking incidents.
Olscore[hidden]
taocoyote[hidden]
Thank you for coding something that I intended to code but have been too lazy to do.
smashed[hidden]
Work is first defined in new or updated specifications, then the change gets made. You have to resist making natural language Todo lists and have the agent write runnable unit tests. My project is a CLI tool so it's fairly easy.
Verifying if it's done is a matter of running the specifications test suite.
In a totally not original way, I'm testing this workflow in my own agent orchestration tool. I know, there's just so many already. I'm building it for myself.
So far, the system is holding up but it's way too early to declare it a success.
You can check it out here:
https://github.com/egzo-ai/egzo/tree/main/specs
I think the same idea could be applied to other projects in different domains.
With this workflow I can read and update the specs, and run it against a built binary and be much more confident that the code works. Claude code has picked up the system without complaining for the most part.
mgw[hidden]
We also open-sourced a similar system called https://github.com/madeinorbit/podium
We do have what you call Items and Shouts in the form of an agent communication system and a Linear style issue tracker that agents and humans use together.
In our case everything around orchestration is discussed with the agent itself and they take action for the user via our CLI. We found any UI to prescribe how teams of agents should coordinate too clunky. So in our case you just tell the agent: "every codex luna max implementer gets an opencode muse spark reviewer. when a subtree is done, review with opus high.". It then sets up the graph and the system enforces the rules.
It works really well for us.
iamjustrach[hidden]
gerwim[hidden]
How do you handle the context? As in: you have a forum like system as you say, how are agents finding the correct "posts"? How do you handle drift or specs that change over time?