Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Bonsai

What becomes possible once a company's context is machine-readable?

Reddit is the place I used to test that question.


Where this started

I built a VSCode extension. Once it was done, I found a harder problem than building it: getting it in front of the right people.

Doing that well on Reddit meant hours of reading subreddits, just to understand what people accept there. I had to track which conversations were relevant, think about the right timing, and check if a comment would sound helpful or like spam. Writing the reply itself was fast. The research and the planning before it took real time, every week, for every community.

That repeated work is what made me think about automating the process behind posting, not just the posting itself.


What already existed

Before this project, I had a data extraction pipeline. It crawls websites, reads both static and dynamic pages, and uses an AI model to pull structured information from a page instead of relying on fixed rules like CSS selectors. It also cleans and stores the data, and exposes the whole process through an API. The code lives in /engine, with its own README.

That code already has its own logic. It's the base I built the rest of this project on, and there's a lot more that can be built on top of it. That's what the rest of this document works through.


The leap

Extraction turns a page into structured data. That's useful on its own, but it made me ask a bigger question: what if that context about a company could stay around, get compared, and get updated over time, instead of being rebuilt for every new task?

That question is where Bonsai started. I used Reddit to test it, because it's a place where a company's problem shows up in the words real people use, not the words the company uses about itself.

                    EXISTING
              AI EXTRACTION LAYER
                       │
                       ↓
              STRUCTURED CONTEXT
                       │
                       ↓
                SEMANTIC MODEL
                       │
          ┌────────────┼────────────┐
          ↓            ↓            ↓
       Research     Monitoring    Strategy
          │            │            │
          └────────────┼────────────┘
                       ↓
                   Validation
                       ↓
                    Feedback
                       │
                       └──────→ updated context

Bonsai is the design I built around that idea: a system that decides, with evidence, if, where, and how a company should join a conversation that's already happening.

The extractor is the part that already works. Everything from here on is the plan I built on top of it. The layers below explain how each piece is meant to work, using the screens as a way to think through the design before building it.


Walking through the design

Company context

It starts with a URL and a short line about what the company wants right now: launch a product, enter a new market, reach a new audience. Bonsai reads the website like a person would, and pulls out what the company sells, who it's for, and what problem it solves. That gets combined with the stated goal. This first step is just raw material, collected but not organized yet.

Input

Semantic understanding

That raw material turns into something structured: a semantic bank. A simple list of keywords only catches pages that use the exact words the company already uses about itself. Anything phrased another way slips past it, and the link between related ideas gets lost too. A keyword list treats "automation" and "too many manual tasks" like two separate things, even though they describe the same problem from two sides. The semantic bank is built to keep those links: products connect to the problems they solve, problems connect to the people who have them, and the system also builds new combinations of words that describe the same idea in different ways. A company that sells "workflow automation" ends up linked to phrases like "too many manual tasks" or "tools that don't talk to each other," even if the website never used those words. That bank, not a list of terms, is what everything after this step gets compared to.

Semantic understanding

Market conversations

Once the semantic bank is ready, the plan is to scan Reddit on an ongoing basis, going through subreddits and threads to find where the ideas in that bank show up, directly or in a related form. This step is meant to be broad. The goal is to find every place where the company's world overlaps with what people are actually talking about, before narrowing anything down.

Opportunity radar

Community context

Every subreddit found in that scan gets its own profile: how technical the discussion tends to be, how it reacts to anything that looks like promotion, what kind of tone gets upvoted or ignored. A community with 2 million members and one with 50,000 can behave very differently, and that profile is what decides how, or if, to approach each one later.

Opportunity detection

Each conversation would get a score. A thread could get picked up just because it has the right words, but that alone says very little, someone could be mentioning the topic in passing, or the thread could already be dead. So the plan is to check how close the conversation's problem is to the company's, if the people in that thread are the right audience, and if the timing is good, a fresh, active thread is worth more than an old one nobody reads anymore. The result would be a ranked list, with the strongest opportunities at the top instead of buried among hundreds of loose matches.

Conversation analysis

Strategic intervention

Once a conversation looks worth entering, the system works out what kind of contribution fits: a comment with a concrete detail, a full post that opens a new discussion, or just watching it for now. Tone can also be adjusted here, a draft can lean more technical, more direct, or more playful, depending on the strategy for that company.

Validation

Before anything gets close to publishing, the plan is to check the draft against the conversation, the community profile, and the company's own context. It would get a score for how well it fits, how likely it is to read as promotion, and how credible it sounds, along with a plain explanation of what works and what to adjust. A person would review the final decision, and this check is what they'd review it against.

Response risk check


Where most of the design work went

Once the system understands the company, the community, and the conversation, writing a plausible reply is the easy part. Deciding if a reply should exist at all is the harder half, and that's where most of the design time went:

  • Should we participate at all?
  • Why this conversation?
  • Why now, and not next week?
  • Why this community, and not a similar one next door?
  • What kind of contribution fits here?
  • What should we avoid saying?

Most of Bonsai's design exists to answer those six questions with evidence, before writing a single word of content.


The discovery that mattered most

A company usually describes its problem in its own words, something like "agent costs." Real conversations often use different words for the same thing: cost per task, execution-path cost, retry budgets, multi-agent spend, language that is new to the company.

Feeding that language back into the semantic model is more than a research result. It's a correction to the model itself.

COMPANY
   ↓
SEMANTIC MODEL
   ↓
REDDIT
   ↓
NEW MARKET LANGUAGE
   ↓
UPDATED SEMANTIC MODEL
   ↓
BETTER RESEARCH
   ↓
BETTER OPPORTUNITIES

If that loop closes, the system could become something bigger than a Reddit tool. It would get closer to a context engine that learns from the market it watches. The same loop is meant to apply to strategy too: what worked in a subreddit, what kind of approach tends to work for a goal, and what has worked for a company before, all of that is meant to feed back into the next recommendation.

Live monitoring

Strategy report


Why this goes beyond Reddit

The Reddit use case is narrow on purpose. The design behind it is meant to work in other places too.

What matters here is the separation between extraction, semantic context, evidence, decision, validation, and feedback. Reddit is one place where that separation is useful. The same design could support any system that needs to understand several companies or clients at once, each with their own context, instead of starting from a blank prompt every time.


What I'd build next

  • Feed real post results, upvotes, removals, replies, back into the scoring
  • Read each community's actual rules, wiki pages and AutoMod settings, instead of only guessing culture from patterns
  • Connect Reddit activity to real results, traffic, mentions, conversions, to close the loop from conversation to outcome
  • Keep publishing as something a person approves, not an automatic action, even as the system grows

Questions Bonsai is designed to answer

  • What does this company actually know how to solve?
  • Where are those problems being discussed right now?
  • Which communities actually matter?
  • What language is showing up around the problem that the company hasn't caught yet?
  • Which conversations are worth entering?
  • What kind of participation fits a community's culture?
  • What should we avoid saying, and where?
  • What has worked before, for this company or a similar one?
  • What should we keep watching next?

Repository structure

bonsai-project/
├── README.md
├── assets/
│   ├── 01-input-url-intent.png
│   ├── 02-semantic-understanding.png
│   ├── 03-opportunity-radar.png
│   ├── 04-conversation-analysis.png
│   ├── 05-response-risk-check.png
│   ├── 06-live-monitoring.png
│   └── 07-strategy-report.png
└── engine/              ← the extraction pipeline this project is built on
    ├── README.md
    ├── ai/
    ├── api/
    ├── crawler/
    ├── fetcher/
    ├── parser/
    ├── pipeline/
    └── storage/

About

AI-powered web data extraction platform with modular pipeline, semantic parsing, and automated data processing using LLMs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages