An AI assistant with a lab notebook no one can quietly rewrite

When an artificial intelligence system helps a scientist write a paper or run an analysis, a basic question follows: how do we know the AI actually did what it claims to have done? A new open-source project addresses that question with an unusually direct answer. It keeps a permanent, tamper-resistant record of the work, stored in an ordinary folder on the researcher's own computer, that the AI itself cannot rewrite.

The project is called K-Dense BYOK, short for "bring your own keys." It is described in a 38-page paper posted to arXiv, the open research preprint service, under the title "K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook." The paper lists four authors, Aubrey M. Brueckner, Darshil Patel, Yuhuan He and Timothy Kassis, and carries arXiv identifier 2610.00074 in the artificial intelligence category.

What the tool is: bring your own keys

K-Dense BYOK (bring your own keys) is a free, open-source AI research assistant for scientists in any field that runs on the researcher's own computer.

The authors, in the paper's abstract on arXiv, submitted September 4, 2026. Source: arXiv abstract page for arXiv:2610.00074, abstract text.

The name describes its economic model. The researcher supplies access to a language model of their choice, either a hosted commercial model or one running locally on their own hardware. The application supplies the rest. According to the abstract, the researcher brings the model and the application brings "a place for the work to run, a layer of scientific scaffolding, and a complete record."

One practical note before going further: according to the project's README on its GitHub repository, K-Dense BYOK is currently in beta. That means the software is still under active development, and readers evaluating it should expect that status to change as the project matures.

That structure matters for access. A scientist who cannot or does not want to pay for a managed cloud research platform can still use a capable model through this kind of tool, on their own machine, under their own control. The code is released under the MIT license, one of the most permissive open-source licenses, which allows anyone to use, modify and redistribute the software, including for commercial purposes.

A folder you own, and a notebook that only grows

Each project lives in what the authors call an ordinary folder. The data, the code, the results and the record all sit in a directory the researcher administers directly.

Each project is an ordinary folder, so the data, the code, the results, and the record stay on a machine the researcher administers and can be read years later without the application.

The authors, in the paper's abstract on arXiv, submitted September 4, 2026. Source: arXiv abstract page for arXiv:2610.00074, abstract text.

This is a meaningful contrast with managed platforms, where the record of work lives on someone else's servers in a format tied to that vendor's product. Here, the record is plain material on the scientist's own disk. If the application disappears, the folder and everything in it remain readable.

The second element is what the authors call a Living Lab Notebook. Its entries, they write, "link into an argument" and "are added to but never erased." In practical terms, this is a research journal that grows over time. Entries can be appended, but the history cannot be quietly edited after the fact, which is what makes the notebook trustworthy as evidence of what happened and when.

Watching, not asking: how the system avoids the AI grading its own homework

The most distinctive design choice is the third: the system does not take the AI's word for what it did.

It records what happened by watching what the agent does rather than by taking the agent's word for it, in a log the agent has no tool that can write to.

The authors, in the paper's abstract on arXiv, submitted September 4, 2026. Source: arXiv abstract page for arXiv:2610.00074, abstract text.

AI research agents are language models given tools: they can search, write files, run code. That also means they can, in principle, describe their own actions inaccurately. A model might claim it verified a result when it did not. The authors call this "model overclaiming" and connect it to their earlier benchmark of nine frontier models. Their design answer, they write, is "making claims checkable rather than preventing them."

The paper's title adds another technical detail: the notebook is hash-chained. Each entry carries a cryptographic fingerprint that depends on the previous entries, so altering an old entry would break the chain. This is the same idea behind tamper-evident logs in other security systems. The agent has no tool that can write to the observed log, so the record of what the agent did is produced by something other than the agent's own account of itself.

For a general reader, the significance is this: the question "how do we know the AI did what it claims?" is answered not by trusting the AI's honesty, but by making the record of its actions independent of the AI's own statements, and readable by anyone, years later.

The benchmark: a self-reported win, stated plainly

The paper includes an evaluation the authors ran themselves. On twenty interdisciplinary research prompts, scored under a rubric fixed in advance, the authors report that K-Dense BYOK led two managed platforms on both scientific quality and research execution.

On twenty interdisciplinary research prompts, scored under a rubric fixed in advance, K-Dense BYOK led two managed platforms on both scientific quality and research execution. Its deliverables were the only ones that recorded the software they ran in, and the only ones that usually arrived with a command that regenerates the results.

The authors, in the paper's abstract on arXiv, submitted September 4, 2026. Source: arXiv abstract page for arXiv:2610.00074, abstract text.

Two details stand out. First, the authors note that one of the managed platforms ran the same frontier model and supplied neither the software record nor the regeneration command, which suggests the difference lies in the scaffolding rather than the raw model. Second, the deliverables' habit of arriving with a command that regenerates the results is a reproducibility feature: another researcher can, in principle, rerun the work and check it.

It is important to be clear about the status of this evidence. This is the authors' self-reported evaluation of their own tool. It is not an independent replication. The scoring rubric was fixed in advance, which reduces one source of bias, and the paper reportedly includes the benchmark prompts, the scoring rubric and per-prompt scores, which allows others to inspect the method. But until outside researchers examine or repeat the comparison, the performance claim should be read as the authors' own finding, not settled fact.

The limitation the authors put in the abstract, not a footnote

The authors do not bury their main limitation. In the abstract they state it directly.

Those environment records were files the agent wrote, not part of the observed log, which does not yet capture the software environment itself.

The authors, in the paper's abstract on arXiv, submitted September 4, 2026. Source: arXiv abstract page for arXiv:2610.00074, abstract text.

In other words, the records of which software ran were produced by the AI agent itself, writing ordinary files. They were not captured by the tamper-resistant observed log, which does not yet record the software environment. That means the strongest protection, the log the agent cannot touch, does not currently cover this particular kind of record. A future version that captured the environment in the observed log would close that gap.

The authors deserve credit for stating this plainly in the abstract rather than in a footnote. It is the difference between claiming a fully tamper-proof research record and honestly describing which parts are protected and which are not. Readers evaluating the tool should keep the distinction in mind: the notebook's append-only design and the agent-unwritable log are design features of the software's architecture as described by its authors; the software-environment record is currently outside that protection.

Why this matters beyond one project

The K-Dense BYOK paper lands in a broader conversation about what happens when AI systems participate in science. If AI increasingly helps generate hypotheses, run analyses and draft papers, the integrity of the research record becomes a question about AI infrastructure, not just about human conduct.

This project frames that question around access and human agency. A free, MIT-licensed tool that runs on the researcher's own computer, works with the model of their choice, and keeps a record the AI cannot rewrite is an argument that trustworthy AI-assisted science does not have to run through a managed platform. Whether it outperforms those platforms in practice is, for now, the authors' claim to be tested. What is verifiable today is that the software exists, the code is openly licensed, the paper and its artifacts are public, and the design makes a specific, checkable bet: that science is better served by a record that cannot be quietly rewritten than by an AI that simply asserts what it did.