Hawker News

Show HN: Terse, a Claude Code plugin that halves reply length by cutting filler

github.com/lowenbjer

24 pointsby lowenbjer18 comments

I struggle with the way Claude speaks to me when I do work. So much text, every, single, time.

When I ask a question, I get an essay back, and with so many weird constructs: slogans, metaphors, recaps, strange choice of nouns, "it's not X, it's Y".

Claude's own "concise" mode shortens it, somewhat, but keeps the the rest of the slop.

I tried system prompts, those got forgotten after a couple of turns. I tried looking for plugins but none consistently made Claude speak normally. Some where slash commands, other solved for token use, others solved for neurodivergence.

I just wanted to not read an essay and make Claude get to the point fast, every single time.

Is that so much to ask? Apparently; it took me weeks of trial and error to make something that worked for me, I have stories, but I'm sharing it as a claude code plugin with you, those of you who struggle with Claude, the same way I did.

TZubiri[6 comments hidden]
Read up on Chain of Thought

The model is essentially thinking out loud, when you ask it to be more concise, you make it think less, therefore producing more erroneous answers.

Some models have an internal chain of thought (claude being one of them), which sometimes isn't even published to avoid reverse engineering, but it seems that this might still be a problem.

What you'd want actually is a layer that summarizes the actual answer, but that's actually an internal prompt by claude that you are not seeing, the model just doesn't expose the necessary bits for you to hack this together.

Try another model that exposes the raw llm output instead of exposing a CoT result directly.

Of course the real hack is learning to read diagonally without reading every single word, this is a skill that is useful in general. It's also less effort in general, instead of making plugins and super customizing the thing, you just consume the default settings, which are hyperoptimized, and require no time spent in configuration.

cyanydeez[4 comments hidden]
I think you're humanizing too much.

The <think> blocks are an attempt to explore the gradient descent space to escape local minimums and find a better global minimum to continue the descent.

While verbosity _might_ do this better, you could easily consider things like "but wait am I forgetting ...." as just one token. So if you actually do it right, you could replace all that with a "hold on" or something of a terse variety.

lowenbjer[3 comments hidden]
I too think there is a fallacy in automatically attributing longer texts to hold more information.
TZubiri[2 comments hidden]
I'm saying quite the opposite, I'm saying that the longer texts are produced not only because they hold more information, but because they are necessary in producing the correct information.

A simple example, if you have two models Short and Long, and ask "what is 1942848x41482982", the Long model might answer in 300 tokens by expanding the operation and producing intermediate factorizations and solutions. While the short answer will simply produce the output. However since the Long model actually developed its answer instead of hallucinating it, it will more often be correct.

It's not that more text holds more information, but that more compute can solve more problems correctly. Limiting output length in models is effectively reducing compute, which obviously will lead to less correct answers.

lowenbjer[hidden]
Ah, nice analogy! The way you describe it makes sense, I need to read up on the topic. Chain of Thought, is that what I should look for?
lowenbjer[hidden]
While i understand your reasoning here, and while I can't produce any proof that that supports my claim that it doesn't, I still would like to share my experiences after using terse personally for a couple of weeks, and those are that I have not noticed the degradation in output quality that you are describing. Quite the opposite, as my interactions use up less context space, my guess, i find that accuracy has increased rather than decreased.
glimshe[2 comments hidden]
My earlier whole grail for LLM interaction was making them stop trying to follow up with a stupid question.

"Are you thinking of doing this or are you just daydreaming?" and garbage like that. They got better at it with LLM options and instructions, but I couldn't make them stop completely to try to keep me "engaged".

lowenbjer[hidden]
I agree, how i use it i don't want it to be creative or help me think unless i ask it to do so. I'm using it to DO work.
Octokat[2 comments hidden]
Doesn't using only custom Output styles work? I've been experimenting with different styles and stripped conversations rules in Claude.md and rules to a bare minimum and that seems to be doing the trick.
lowenbjer[hidden]
Yeah i tried that but I did not get it to work consistently.

It feels like it took the styles as suggestions rather than rules to enforce.

How has it been working out for you in longer conversations?

trashymctrash[2 comments hidden]
Opus 5.5 has reduced this to an acceptable level for me. Are you also using this plugin with that model?
lowenbjer[hidden]
Yes! Some of my benchmarks was against Opus 5.5, and it reduced output over 50% on vanilla settings. Check repo for data.
preetam960[2 comments hidden]
You said system prompts got forgotten after a couple of turns. How does Terse avoid that? Is the instruction re-added every turn from a UserPromptSubmit hook, or something else?
lowenbjer[hidden]
Yes, exactly that, so there's refined writing instructions in the system prompt and a reminder with a compact version of the rules injected every turn from a UserPromptSubmit hook.
citizenfishy[2 comments hidden]
Claude Mods seem to answer this pain better without affecting the model reasoning
lowenbjer[hidden]
How so?
userbarznji[hidden]
asma tuana
userbarznji[hidden]
hellophoto