Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .cursor/rules/use-local-code-intel.mdc
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,8 @@ MCP namespace: `user-local-code-intelligence`. The corpus is already embedded in

## Always do this first

1. `search_codebase` (pass `max_tokens` 800–1500). Only the query is embedded locally; stored vectors retrieve snippets.
2. `search_symbol` for a known name; `find_references` for occurrences.
1. For a new or broad coding task, `get_task_context`. Only the query is embedded locally; stored vectors retrieve snippets.
2. `search_symbol` for a known name; `search_codebase` for conceptual exploration; `find_references` for occurrences.
3. `get_file_context` with `start_line`/`end_line` for the range you will edit.
4. `get_repo_context` / `list_indexed_repos` / `index_status` for orientation.

Expand All @@ -20,4 +20,4 @@ MCP namespace: `user-local-code-intelligence`. The corpus is already embedded in
- Read whole files to "see how it works"
- Ask the cloud model to embed or index code

If the workspace is a parent (e.g. Savor), `list_indexed_repos` and pass `repo`. If unindexed, run `code-intel setup --repo <path>` via Shell, then search again.
If retrieval reports low confidence, the index is stale, or the path is unindexed, targeted filesystem search is allowed. If the workspace is a parent (e.g. Savor), `list_indexed_repos` and pass `repo`. If unindexed, run `code-intel setup --repo <path>` via Shell, then search again.
3 changes: 3 additions & 0 deletions .github/PULL_REQUEST_TEMPLATE.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,3 +7,6 @@
- [ ] `npm run typecheck`
- [ ] `npm run test:unit`
- [ ] (optional) `npm test` with Ollama + `nomic-embed-text` if you touched indexing or MCP search
- [ ] Retrieval changes include before/after benchmark evidence
- [ ] CLI, MCP, privacy, or storage changes include documentation updates
- [ ] No credentials, private source, organization endpoints, or local index data are included
53 changes: 53 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# Changelog

All notable changes to this project are documented here. The project follows [Semantic Versioning](https://semver.org).

## [Unreleased]

## [0.2.0] - 2026-09-21

### Added

- `get_task_context`, a high-level MCP tool that returns a ranked, token-budgeted context package for a coding task.
- `code-intel context`, retrieval modes, confidence reporting, and `--explain` traces.
- Deterministic task intent, relationship expansion, diversity controls, and hard context budgets.
- Exact symbol and basename ranking floors, score-ordered files, and test/config intent handling.
- Lightweight chunk metadata for imports, exports, referenced identifiers, tests, and configuration files.
- TypeScript path-alias and Python module expansion.
- Labeled retrieval benchmark with Precision@5, coverage@5, Recall@10, MRR, NDCG, token reduction, latency, and auditable retrieved paths.
- Per-file watcher updates for small event batches, content-sample freshness detection, and persisted watcher errors in `index_status`.
- Actionable CLI and MCP recovery guidance plus confidence-based targeted filesystem fallback.

### Changed

- Cursor rules, skill, hooks, and MCP instructions now prefer `get_task_context` for broad tasks.
- Hybrid retrieval now combines semantic, keyword, symbol, path, structure, relationship, test, and recency signals.
- MCP starts an incremental catch-up when it opens an indexed workspace.
- Documentation now separates measured benchmark claims from estimates and describes the privacy boundary for each provider.

### Fixed

- Duplicate symbol chunk IDs no longer abort a LanceDB merge.
- Unchanged chunks reuse embeddings while refreshing metadata and line ranges.
- Large event bursts fall back to full incremental discovery instead of issuing many individual writes.
- Code-change queries that rank documentation first are treated as low confidence.

## [0.1.1] - 2026-09-07

### Added

- Guided onboarding, Ollama health checks, and Cursor integration.
- OpenAI-compatible embedding provider support and Git-aware multi-repository indexing.
- Parallel file indexing and IVF-PQ vector indexing for larger tables.

## [0.1.0] - 2026-09-01

### Added

- Initial public npm release.
- Structural chunking, local LanceDB storage, incremental indexing, hybrid search, MCP tools, and file watching.

[Unreleased]: https://github.com/pranitmodi/code-intel/compare/v0.2.0...HEAD
[0.2.0]: https://github.com/pranitmodi/code-intel/compare/v0.1.1...v0.2.0
[0.1.1]: https://github.com/pranitmodi/code-intel/releases/tag/v0.1.1
[0.1.0]: https://github.com/pranitmodi/code-intel/releases/tag/v0.1.0
2 changes: 2 additions & 0 deletions CODE_INTEL_NEXT_PHASE.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,7 @@
# code-intel: Next-Generation Agent Context Architecture

> Historical implementation specification for the task-context work introduced in v0.2.0. Some acceptance items are now implemented and some remain future work. See the maintained [architecture](docs/ARCHITECTURE.md), [benchmarks](docs/BENCHMARKS.md), and [changelog](CHANGELOG.md) for current behavior.

## Purpose

This document is the implementation specification for evolving `code-intel` from a local semantic code-search server into a **persistent context layer for AI coding agents**.
Expand Down
41 changes: 41 additions & 0 deletions CODE_OF_CONDUCT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
# Code of Conduct

## Our pledge

We pledge to make participation in the `code-intel` community a harassment-free experience for everyone, regardless of age, body size, disability, ethnicity, sex characteristics, gender identity and expression, experience level, education, socioeconomic status, nationality, personal appearance, race, caste, color, religion, sexual identity and orientation, or technology choices.

We pledge to act and interact in ways that contribute to an open, welcoming, diverse, inclusive, and healthy community.

## Expected behavior

Examples of positive behavior include:

- showing empathy and respect;
- giving and accepting constructive technical feedback;
- focusing discussion on evidence, reproducible behavior, and project goals;
- respecting privacy, especially when reports involve private source or credentials;
- accepting responsibility and learning from mistakes.

Unacceptable behavior includes:

- harassment, threats, insults, or discriminatory language;
- trolling, sustained disruption, or personal attacks;
- publishing another person's private information without permission;
- posting private source, credentials, or sensitive logs;
- conduct that would reasonably be considered inappropriate in a professional setting.

## Enforcement

Project maintainers are responsible for clarifying and enforcing these standards. They may edit, remove, or reject comments, commits, code, issues, and other contributions that violate this Code of Conduct, and may temporarily or permanently ban contributors for behavior they consider harmful.

Report conduct concerns privately to the repository owner through the contact options on the [GitHub profile](https://github.com/pranitmodi). Do not use a public issue for a sensitive report.

All reports will be reviewed promptly and fairly. Maintainers will respect the privacy and security of reporters.

## Scope

This Code of Conduct applies in project spaces and when an individual is officially representing the project in public spaces.

## Attribution

This Code of Conduct is adapted from the [Contributor Covenant](https://www.contributor-covenant.org), version 2.1.
34 changes: 32 additions & 2 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@

Thanks for helping. The goal of this project is a **local-first** code index: embeddings and source stay on the contributor's machine, and AI agents query that index over MCP instead of re-scanning the tree.

Before changing a subsystem, read [Architecture](docs/ARCHITECTURE.md). Retrieval changes must follow the measurement rules in [Benchmarks](docs/BENCHMARKS.md). Participation is governed by the [Code of Conduct](CODE_OF_CONDUCT.md).

## Prerequisites

- Node.js 20+
Expand Down Expand Up @@ -30,6 +32,7 @@ Unit tests in `tests/unit` never need Ollama. Integration tests in `tests/integr
```bash
npm run typecheck
npm run test:unit # CI default
npm run benchmark:retrieval # labeled retrieval metrics (needs an index + embeddings)
npm test # unit + integration (integration skipped without Ollama)
npm run dev -- status # CLI from source, no build step
```
Expand All @@ -50,24 +53,51 @@ Indexes live under `~/.local-code-intelligence` (or `CODE_INTEL_DB_PATH`), not i

`tests/integration/mcp.test.ts` spawns `code-intel mcp` over stdio with `@modelcontextprotocol/client` — the same transport Cursor uses. When you add or rename a tool, update that file's expected tool list.

## Retrieval and benchmark changes

Retrieval quality is part of the product contract. A ranking change should include:

- focused unit tests for the intended signal;
- `code-intel search "<query>" --explain` or `code-intel context "<task>" --explain` output for the affected case;
- before/after `npm run benchmark:retrieval` results;
- an explanation for any task-label or relevance-grade change.

Do not improve metrics by broadening the context until it resembles a tree scan. The release floor is relevant-file coverage@5 ≥ 0.60, Recall@10 ≥ 0.75, MRR ≥ 0.70, and average per-task token reduction ≥ 60% on the repository dataset.

When adding a benchmark task, use a realistic engineering request and defensible relevant files. Avoid tasks designed around the current ranker's implementation.

## Publishing (maintainers)

Releases are made from a clean, tested `main` commit:

```bash
npm run typecheck
npm run test:unit
npm run build
npm run benchmark:retrieval
npm audit --omit=dev
npm pack --dry-run
npm login
npm publish --access public
```

The package name is `@pranitmodi/code-intel` (scoped; unscoped `code-intel` is blocked by npm as too similar to `codeintel`).

`prepublishOnly` builds `dist/` and runs unit tests. The tarball includes `dist/` plus README and LICENSE; indexes under `~/.local-code-intelligence` are never published.
`prepublishOnly` builds `dist/` and runs unit tests. Inspect the tarball before publishing. It must contain only the built CLI/server and public documentation—never indexes, `.code-intel` config, credentials, private paths, benchmark secrets, or `~/.local-code-intelligence` data.

Update [CHANGELOG.md](CHANGELOG.md), bump `package.json` and `package-lock.json` together, tag the exact published commit, verify it with `npm view`, and create a matching GitHub release.

## Pull requests

- Keep changes focused; match the existing module boundaries (`src/discovery`, `src/chunker`, `src/embeddings`, `src/vector-store`, `src/indexer`, `src/search`, `src/mcp`).
- Keep changes focused; match the existing module boundaries (`src/discovery`, `src/chunker`, `src/embeddings`, `src/vector-store`, `src/indexer`, `src/search`, `src/retrieval`, `src/benchmark`, `src/mcp`).
- Do not commit `node_modules/`, `dist/`, or anything under `~/.local-code-intelligence`.
- Do not add cloud embedding APIs as the default path; local Ollama is the contract.
- Run `npm run typecheck` and `npm run test:unit` before opening a PR.
- Keep public examples provider-neutral. Do not commit organization-specific endpoints, usernames, paths, source, or benchmark credentials.
- Update documentation when changing CLI commands, MCP tools, privacy boundaries, storage, or watcher behavior.

## Security

Secret-shaped files (`.env*`, keys, credentials) are excluded from indexing by default. Do not weaken those filters without a documented, opt-in config flag.

Do not report vulnerabilities or credential exposure in a public issue. Follow [SECURITY.md](SECURITY.md) and use GitHub private vulnerability reporting.
Loading
Loading