Ed's Lab Targetron Knowledge Catalog

Working project / September 2026

Targetron Knowledge Catalog

I wanted to see whether a knowledge-catalog workflow I had already worked through for Outscraper could be reused for another real website instead of rebuilding the process from scratch.

The Targetron pilot gave me a practical way to test URL discovery, routing, crawling, QA, classification, and refresh logic while also forcing me to document where a reusable workflow needs clearer rules.

Back to the Lab → Repository: Private
Status Real-site pilot complete

The reusable workflow and interface work are still evolving.

Site tested Targetron.com

Public website discovery and knowledge-catalog workflow.

Main tools Python · GitHub · YAML

With CSV/JSONL artifacts, PowerShell, VS Code, and validation steps.

Repository note: The source repository is private because this project is tied to company-specific work. I am sharing the process, results, and lessons publicly without exposing the private codebase or internal artifacts.

Reuse the workflow on a second real website.

I started with a reusable knowledge-catalog structure rather than a blank repository. The goal was to point the workflow at Targetron, keep the brand-specific rules separate, and see how much of the process could stay reusable.

The workflow covered site discovery, URL routing, metadata crawling, QA, taxonomy and classification work, catalog generation, graph and link intelligence, coverage checks, overlap or refresh signals, and a refresh baseline.

01 Discover

Collect the public URL inventory from the site's sitemap sources.

02 Route

Separate semantic content from reference or navigation-oriented URLs.

03 Crawl + QA

Collect metadata and send uncertain records through review gates.

04 Classify

Apply controlled taxonomy and keep approved knowledge separate from generated intelligence.

I wanted something more useful than keeping a large site in my head.

Content work across a growing SaaS site involves product pages, articles, categories, reference URLs, internal links, and older content that can overlap with newer pages. A structured catalog gives me a way to inspect that system instead of relying on memory.

The second reason was technical. I wanted to know whether the workflow was actually reusable. A process can look reusable when it has only been tested once. Applying it to Targetron exposed the parts that belonged in the common framework and the parts that needed brand-specific configuration.

The pilot made it through the full workflow.

The discovery and routing stages produced a usable inventory, and the later QA and classification stages reached a clean review state.

798 unique URLs discovered
698 semantic / navigation URLs routed
100 reference URLs routed separately
153 candidates crawled
0 unresolved records at the final gate

The final gate contained 133 semantic records and 20 reference-only records, with no unresolved records left at that stage. I also ran a 20-page classification pilot and a 20-page refresh baseline.

Discovery was easier than deciding what each URL should mean.

The first challenge was that a discovered URL is not automatically useful knowledge. Targetron's discovery run found 798 URLs, but the workflow still had to separate 100 business-data reference URLs from the larger semantic and navigation inventory.

One sitemap source also failed while eight succeeded. That reinforced an important lesson for me: discovery needs its own reporting and QA. I should not assume every candidate source works just because the final URL list looks plausible.

I also ran into a practical project-organization problem while moving between the reusable template, the Targetron-specific repository, and the separate web interface. Similar folder and repository names made it easy to run a command from the wrong location. That sounds small, but it showed me why setup instructions and repository responsibilities need to be explicit.

Reusable workflows need boundaries, not just reusable code.

The project helped me understand that discovery, routing, crawling, human review, taxonomy, classification, and derived intelligence are different jobs. Keeping them separate makes the system easier to inspect and easier to explain.

It also changed how I think about automation. Rules and AI can propose classifications, relationships, and review signals, but I do not want them silently approving changes to the knowledge base. Controlled taxonomy, review gates, and regenerable outputs make the workflow easier to trust.

From a GitHub perspective, I am also learning that a useful technical project is more than the Python scripts. The repository structure, configuration files, generated artifacts, validation commands, and README all affect whether I can understand the project again later.

This sits between SEO, content operations, and technical writing.

SEO

A structured URL inventory can support content-gap review, overlap checks, internal-link analysis, and refresh decisions.

Content operations

The catalog gives me a repeatable reference for understanding what already exists before deciding what should be created or updated.

Technical writing

I have to document the workflow well enough that I can tell which steps are reusable, which are brand-specific, and where human review belongs.

APIs and automation

The next layer connects background jobs, persisted artifacts, API results, and browser-based review panels instead of treating each script as an isolated task.

Make the workflow easier to inspect from the browser.

The real-site pilot proved that the workflow can run end to end. My next work is less about adding another crawler and more about making the workflow easier to operate and review.

  1. Continue connecting background jobs to API results and persisted artifacts.
  2. Improve the browser panels for discovery, routing, metadata, and QA review.
  3. Keep the reusable framework separate from Targetron-specific configuration.
  4. Document the workflow well enough that I can repeat the process on another site.

The useful part of this project is not that I built a crawler. It is that I am learning how discovery, classification, review, documentation, and maintenance fit together as one content-data workflow.

More experiments → Repository: Private