I created a simple git repository with company skills. Basically just a collection of skills around tools and practices we share. One of the skills is "update company skills" this simply pulls the changes from git and wires them into the user's ~/.codex directory. You can probably do something similar for claude code.
This is far from perfect but we're in this weird transition phase where none of the major AI tool providers are really focusing much on team use of their stuff. But I expect that will start changing soon.
Current tools mostly focus on individuals doing things in isolation. And of course in a team there's more to collaborating than throwing stuff at each other via github. A central repository of company skills is merely our way of improvising a solution.
I find it interesting that Anthropic hired a few of the key people behind Zulip recently. Team chat with tightly integrated AI tools could be a missing piece here. Team communication flows and processes, including ways of working and guardrails are sort of the next piece of the puzzle here. Going from everyone doing their own thing to teams and companies doing things together is going to be a bit of a journey.
IMO, these agentic guard rails aren't the answer. It seems like we're seeing that the more you stuff context, the more the agents forget and don't follow the guidelines.[1]
Things like ArchUnit, static analyzers, and other deterministic tools can help with lower level things like architecture. For higher up stuff, I am increasingly feeling like agents don't guarantee anything and in many cases its just the opposite. This is where a thoughtful engineer and reviewer can keep things in check.
It's possible I'm off base here, but I can't make heads or tails of the LLM written readme.
I feel like this is going to add 100K tokens and 10K rules to everything that the agent is trying to do. Sort of like dumping the Clean Code book into context and saying, hey now you know how to code cleanly, write great code now
I see what this is going for but can’t help but feel like it’s overwrought. It feels a bit like the most likely outcomes are increased token burn, review surface, and time per task vs not using this.
I feel that way about a lot of very long instructions.
It seems that many of these projects are not benchmarked, so it's difficult to know whether there is an improvement in any circumstance, and what the cost is. Of course, a benchmark will be fuzzy, because codebases are all different, but it'd be a start.
Many of these skills, rules, "playbooks", and such are kitchensink attempts at steering the model. It augments the model to frame it's reasoning according to project rules and flows, but cannot be really trusted to adhere to it. More of a vibe guideline.
Factor III — Mission Definition. Factor IV — Structured Planning. Factor IX — Traceability.”
Why are the Roman numerals in that order, what do they reference and why. Why is this Pilar 2. Etc. It’s easy to get lost in this, even if it would be the best approach in the world.
why is so hard for people to create a CONTRIBUTE.MD file which tells anyone (including AI) how to contribute to the repo. You can also set gates and everything.
I work with an engineering team using Claude Code, Codex, and OpenCode daily. Early on, we created a custom fork of spec-kit to set up our team directives repository and distribute custom agent workflows.
While that fork got us started, maintaining a custom fork of a CLI just to ship prompt workflows created constant merge debt and maintenance headaches. Every session, agents would still drift or forget our architecture decisions, and prompt shortcuts alone couldn't enforce team-wide standards across different developer tools.
To eliminate the fork entirely, we decoupled our workflow skills into adlc-team-skills, built on the open Agent Skills standard (SKILL.md). You install them into any repo with npx skills add tikalk/adlc-team-skills.
The setup works across a few core layers:
On session start, team-boot auto-loads your team constitution from Git and dynamically fetches only the rules, PDRs, and ADRs relevant to the active task — zero prompt-wall bloat.
For product and architecture strategy, product decisions are captured as Product Decision Records (PDRs) and compiled into PRD.md, while architectural decisions use Rozanski and Woods viewpoints composed into AD.md.
For execution, mission-brief acts as an autonomous pipeline runner that derives a formal contract (Goal, Constraints, Non-Goals, Success Criteria) and walks a specify-plan-implement-converge loop. When an agent fails, you edit the spec, not just the code.
In v0.15.0, mission-brief auto-discovers installed skills at runtime. Whether you have spec-kit, OpenSpec, Matt Pocock's skills, Addy Osmani's checklists, or custom skills installed side-by-side, the LLM dynamically decides which skill fits each pipeline step — letting us run upstream spec-kit directly with zero custom fork code.
What doesn't work well yet: our evals suite holdout-split validation is still manual. The architecture skills work, but multi-view DAG orchestration can be slow on very large codebases.
I'm curious — for those of you managing AI coding agents across engineering teams, how are you balancing team standards with the maintenance overhead of custom agent tooling?
I created a simple git repository with company skills. Basically just a collection of skills around tools and practices we share. One of the skills is "update company skills" this simply pulls the changes from git and wires them into the user's ~/.codex directory. You can probably do something similar for claude code.
This is far from perfect but we're in this weird transition phase where none of the major AI tool providers are really focusing much on team use of their stuff. But I expect that will start changing soon.
Current tools mostly focus on individuals doing things in isolation. And of course in a team there's more to collaborating than throwing stuff at each other via github. A central repository of company skills is merely our way of improvising a solution.
I find it interesting that Anthropic hired a few of the key people behind Zulip recently. Team chat with tightly integrated AI tools could be a missing piece here. Team communication flows and processes, including ways of working and guardrails are sort of the next piece of the puzzle here. Going from everyone doing their own thing to teams and companies doing things together is going to be a bit of a journey.
IMO, these agentic guard rails aren't the answer. It seems like we're seeing that the more you stuff context, the more the agents forget and don't follow the guidelines.[1]
Things like ArchUnit, static analyzers, and other deterministic tools can help with lower level things like architecture. For higher up stuff, I am increasingly feeling like agents don't guarantee anything and in many cases its just the opposite. This is where a thoughtful engineer and reviewer can keep things in check.
It's possible I'm off base here, but I can't make heads or tails of the LLM written readme.
1: https://arxiv.org/html/2510.05381v1
see the architect adr and Product pdr, also the whole idea is to have an index of context directive similar to skills which the modal invoke
I feel like this is going to add 100K tokens and 10K rules to everything that the agent is trying to do. Sort of like dumping the Clean Code book into context and saying, hey now you know how to code cleanly, write great code now
Exactly, none of these skills really help other than bloating your token usage
Got rid of superpowers and other useless skills
Your LLM is more than capable of learning from the sea of knowledge
I see what this is going for but can’t help but feel like it’s overwrought. It feels a bit like the most likely outcomes are increased token burn, review surface, and time per task vs not using this.
Who knows, maybe that’s a good thing.
I feel that way about a lot of very long instructions.
It seems that many of these projects are not benchmarked, so it's difficult to know whether there is an improvement in any circumstance, and what the cost is. Of course, a benchmark will be fuzzy, because codebases are all different, but it'd be a start.
Relinking the recent study which argues that many instructions in the context are not followed through https://news.ycombinator.com/item?id=49096969
Many of these skills, rules, "playbooks", and such are kitchensink attempts at steering the model. It augments the model to frame it's reasoning according to project rules and flows, but cannot be really trusted to adhere to it. More of a vibe guideline.
see the team-boot skill, it is lean
The promise is nice, yet I find the readme very hard to understand. So I’m not sure how it solves these problems exactly.
First off I would expect a (team) methodology to be referenced. There are tons to choose from. From that point on other terms may make more sense.
For example:
“ Pillar 2: Product Strategy & Architectural Governance (PDRs & ADRs)
Factor III — Mission Definition. Factor IV — Structured Planning. Factor IX — Traceability.”
Why are the Roman numerals in that order, what do they reference and why. Why is this Pilar 2. Etc. It’s easy to get lost in this, even if it would be the best approach in the world.
I couldn’t get past the first couple lines. It’s clear no human reviewed that.
Thanks for the feedback, working on that
Write your README or ME won't READ it.
Sure, will be fixing that.
why is so hard for people to create a CONTRIBUTE.MD file which tells anyone (including AI) how to contribute to the repo. You can also set gates and everything.
> Stop Vibe Coding in Silos. Build a Shared Cognitive Layer for Your Engineering Team.
I dont want to be biased but i can get myself to read this after an opening like that .
Slop
I work with an engineering team using Claude Code, Codex, and OpenCode daily. Early on, we created a custom fork of spec-kit to set up our team directives repository and distribute custom agent workflows.
While that fork got us started, maintaining a custom fork of a CLI just to ship prompt workflows created constant merge debt and maintenance headaches. Every session, agents would still drift or forget our architecture decisions, and prompt shortcuts alone couldn't enforce team-wide standards across different developer tools.
To eliminate the fork entirely, we decoupled our workflow skills into adlc-team-skills, built on the open Agent Skills standard (SKILL.md). You install them into any repo with npx skills add tikalk/adlc-team-skills.
The setup works across a few core layers:
On session start, team-boot auto-loads your team constitution from Git and dynamically fetches only the rules, PDRs, and ADRs relevant to the active task — zero prompt-wall bloat.
For product and architecture strategy, product decisions are captured as Product Decision Records (PDRs) and compiled into PRD.md, while architectural decisions use Rozanski and Woods viewpoints composed into AD.md.
For execution, mission-brief acts as an autonomous pipeline runner that derives a formal contract (Goal, Constraints, Non-Goals, Success Criteria) and walks a specify-plan-implement-converge loop. When an agent fails, you edit the spec, not just the code.
In v0.15.0, mission-brief auto-discovers installed skills at runtime. Whether you have spec-kit, OpenSpec, Matt Pocock's skills, Addy Osmani's checklists, or custom skills installed side-by-side, the LLM dynamically decides which skill fits each pipeline step — letting us run upstream spec-kit directly with zero custom fork code.
What doesn't work well yet: our evals suite holdout-split validation is still manual. The architecture skills work, but multi-view DAG orchestration can be slow on very large codebases.
Repos: - Skills: https://github.com/tikalk/adlc-team-skills - Methodology: https://github.com/tikalk/agentic-sdlc-12-factors - CLI: https://github.com/tikalk/adlc-skills-cli
I'm curious — for those of you managing AI coding agents across engineering teams, how are you balancing team standards with the maintenance overhead of custom agent tooling?