repo2nb

Turn a GitHub repo into a runnable notebook without the manual work.

Turning a GitHub repo into a runnable notebook usually means wrestling with dependencies and file paths by hand.

4 min readMachine Learning

Every data practitioner has faced the same silent struggle: you clone a paper's repository, open the code, and immediately wonder what you are supposed to run first. The README is sparse, the dependencies are a mystery, and the file tree looks like a puzzle box. repo2nb 0.2.0 steps into that gap with a straightforward promise: turn that unfamiliar GitHub repo into a runnable Kaggle or Colab notebook without you having to reverse-engineer someone else's project structure. The update brings a few thoughtful refinements, and the direction is worth paying attention to, especially if you spend time with open-source code you did not write.

The dependency resolution logic is the quiet workhorse here. It tries poetry export, then uv export, then requirements.txt, and only falls back to an AST import scan when none of those exist. That order makes sense for the modern Python ecosystem, where lockfiles and dependency managers are increasingly common. But the real insight is that the output is always a plain `%pip install` cell, regardless of which path resolved. That means the tool is not locking you into a specific package manager at runtime; it only needs those tools locally during generation. This is a smart design choice because it keeps the generated notebook portable and avoids assuming the execution environment has poetry or uv installed. For a reader who has ever fought with environment drift between their local machine and a cloud notebook, this is a quiet relief.

The reverse mode and incremental sync round out the update in ways that suggest the tool is thinking about the full workflow, not just the one-way conversion. Reverse mode reconstructs the repo from a generated notebook using per-cell metadata, which is a clever use of the hash and path information that is now embedded in every cell. It also validates against directory traversal and refuses to write into a non-empty directory without `--force`, which is the kind of guardrail that shows an awareness of real-world footguns. Incremental sync goes the other direction, keeping the notebook in step with upstream changes to the repo, with a `--dry-run` flag for previewing the diff. These features matter because they turn repo2nb from a one-off conversion script into a maintenance tool, something you can use repeatedly as a project evolves.

The fallback order for dependency resolution is the one open question worth probing. Poetry and uv cover a lot of ground, but what about projects that use conda environments or pipenv? The AST import scan is a decent last resort, but it can miss dynamic imports or optional dependencies. We would tell a reader considering this tool to treat the fallback order as a starting point, not a guarantee. If you regularly work with repos that use less common dependency management, you might still need to hand-edit the generated install cell. That is a minor friction, but it is worth knowing before you rely on this for a critical workflow.

What stands out across all of this is that the tool is designed for people who want to spend less time on setup and more time on the actual code. That aligns with a broader theme we have been tracking, whether it is catching AI slop before it skews your model or exploring real-world computer vision deployments. The common thread is that tooling should remove friction, not add it, and repo2nb is aiming squarely at that goal. The fact that it targets Kaggle and Colab specifically also signals a recognition that cloud notebooks are where a lot of exploratory work happens now, not just in a local IDE. For readers who live in those environments, this could be the difference between actually running a paper's code and just reading it. The specific detail to watch is whether the maintainer extends the dependency resolution to cover conda or pipenv, because that would close the last real gap in an otherwise practical tool.

From Machine Learning

repo2nb is an open-source CLI that converts a GitHub repo into a runnable Kaggle or Colab notebook: walks the file tree, resolves dependencies, and generates cells, instead of you doing that by hand for a repo you didn't write (a paper's code, a tutorial, someone else's experiment).

Repo: https://github.com/David-Magdy/repo2nb

Read the original at Machine Learning