A luminous future knowledge archive representing: Survey the Work Before You Ingest It.

August 21, 2026. I am running VALESKA's discovery method against our own organization before trusting it with a client's documents. VALESKA is the document and organizational-memory system I am building. Ingestion means bringing material into that system for processing and retrieval. Starting there would make classification and scope decisions before I have established what the business actually holds.

The first internal discovery pass records fourteen defects in the method and its supporting tools. Three require attention before proceeding: missing questions about personal and medical material, missing questions about former-client confidential material, and a required inventory tool that does not yet exist at that point in the day's work. I build the inventory later today. Finding the omissions at home is useful; it is not evidence that the method is ready for every client.

Look before collecting

I document a survey ladder that starts with a reachability check, then progresses through structural counts, names and folder patterns, document profiling, content sampling, and human validation of the proposed model. The document calls it five tiers but numbers the preliminary check as Tier 0, followed by five further stages. The practical distinction is between learning from the structure and opening documents. More expensive inspection should answer a question the cheaper survey cannot settle.

Folder names matter because they show how people already organize their work: by customer, project, property, or matter. Those terms are candidates for a controlled vocabulary, the agreed names an assistant can use consistently. They still need confirmation from the business. In one measured collection, 55.4% of 2,019 classified files fell into a single process category that was semantically wrong. I will not treat a standard category list as understanding the company.

The count is a finding

The new dry-run inventory reads without depositing documents or updating a database. It distinguishes raw file counts from plausible document counts and surfaces duplicate content, filename collisions, and sensitivity flags. Hashing creates a fingerprint for comparison; it does not make an inventory correct by itself. Comparing normalized document text can identify equivalent content that a byte-for-byte file comparison misses. Reading that content is also more work than listing names.

Today's evidence shows why I need those distinctions. Windows directory junctions, which point to other directories, inflated one count from 22,281 to 78,679. Counting a whole excluded dependency directory as one item inverted the reported noise proportion from 99.1% to 1.9%. Across four surveyed collections, the scope note reports 405,711 files from a naive walk, 11,671 real files, and 3,700 actual documents. These are different populations, not interchangeable estimates of an ingestion workload.

Scope needs an owner

A sensitivity flag is a question for an authorized person, not an instruction to silently discard a file. Personal health information and prior-client records need explicit scope decisions. Source code is outside the document engagement by default, but prose inside a repository, such as a runbook or architecture note, can remain in scope. The boundary is the asset, not merely the folder containing it. Code requires separate authorization and treatment, not accidental collection alongside documents.

There is a substantial tooling limit in today's survey document: the document profiler has only run on structured material, not Word, Excel, PowerPoint, or PDF. In one surveyed consulting folder, 74% of the files were binary Office documents. A proposed profiling stage is not demonstrated coverage of those formats. Local and network survey rates also differ sharply, so my timings are observations on this equipment, not a client's schedule.

The next evidence

I want a repeatable inventory of one agreed collection, an owner-reviewed record of sensitive and excluded material, and profiling tested on the formats actually present. Only then can I defend the proposed classification model and the work estimate. Today's result is a better survey and a clearer list of missing evidence, not a completed client ingestion.


Historical basis: VALESKA commits 53e486c, c94515e, e642f5a, 34990fa, August 21, 2026 (Pacific time).