Thank you for your interest in contributing to codeloom! This guide will help you get started.
# Clone the repository
git clone https://github.com/algodesigner/codeloom.git
cd codeloom
# Create a virtual environment
python -m venv .venv
source .venv/bin/activate # Linux/macOS
# .venv\Scripts\activate # Windows
# Install in development mode with dev dependencies
pip install -e ".[dev]"# Run all tests with coverage
pytest
# Run a specific test file
pytest tests/test_store.py
# Run with verbose output
pytest -v
# Run property-based tests
pytest tests/test_graph_properties.py -v
# Run stress tests (small only)
SKIP_STRESS=0 pytest tests/stress/ -v
# Run full stress suite (slow — includes 10k-file repo)
FULL_STRESS=1 SKIP_STRESS=0 pytest tests/stress/ -vWe use Ruff for linting and formatting:
# Check for issues
ruff check .
# Auto-fix issues
ruff check --fix .
# Format code
ruff format .Key conventions:
- Line length: 100 characters
- Target Python: 3.10+
- Import sorting: isort-compatible (handled by Ruff)
codeloom/
├── cli/ # Click-based CLI interface
├── core/ # Pipeline stages (detect, extract, build, cluster, analyze)
├── query/ # Hybrid search engine (vector + graph + keyword + RRF)
└── storage/ # SQLite + FAISS storage layer
- Fork the repository and create a feature branch from
main. - Write tests for any new functionality in
tests/. - Run the test suite to ensure nothing is broken.
- Follow the existing code style — Ruff will help enforce this.
- Keep commits focused — one logical change per commit.
- Keep PRs focused on a single change.
- Include a clear description of what the PR does and why.
- Ensure all tests pass before submitting.
- Update documentation if you change public APIs or CLI commands.
The pipeline follows a linear flow:
detect → extract → build → embed → cluster → analyze → store
- detect: Scans directories, classifies files by language.
- extract: Tree-sitter AST extraction with regex fallback.
- build: Assembles a NetworkX DiGraph with deduplication.
- embed: Generates sentence-transformer embeddings locally.
- cluster: Hierarchical Leiden community detection.
- analyze: Structural analysis (god nodes, hubs, quality metrics).
- store: SQLite + FTS5 + FAISS vector index, all in a single file.
- Use GitHub Issues for bug reports and feature requests.
- Include reproduction steps for bugs.
- Mention your Python version and OS.
By contributing, you agree that your contributions will be licensed under the MIT License.