diff --git a/TODO.md b/TODO.md index aa3e976..9f7a277 100644 --- a/TODO.md +++ b/TODO.md @@ -3,6 +3,7 @@ ## At next release cut (v0.18 promote) - [ ] Add `versions/latest/service/rest-api/client/list-embedding-models` to the REST API "Client" nav group in docs.json (page is staged in `next/` only; `scripts/docs.py` will also flag it during promote) +- [ ] C++ docs are removed in `next/`. Before running `promote`, delete `versions/latest/embedded/cpp/`, the C++ nav group for the latest version in docs.json, and point the `/cpp-embedded` redirect to `/versions/latest/embedded/python/introduction`. Until then, `promote` refuses and `sync-next` re-creates the pages in `next/`. ## Should have @@ -35,12 +36,6 @@ - [ ] Document async patterns more thoroughly - [ ] Add version compatibility note (e.g., "Compatible with langchain-core >= 0.2.0") -#### C++ -- [ ] Add thread-safety documentation (if applicable) -- [ ] Add memory management guidance (unique_ptr vs raw pointers) -- [ ] Add performance notes to methods like Train() and Query() -- [ ] Create a "Migration from Python" guide for C++ users - #### Python - [ ] Add note about cache size units (bytes vs megabytes) - clarify in create-index and load-index docs - [ ] Add examples of multi-tenancy use cases in encrypted-indexes/introduction.mdx @@ -50,8 +45,6 @@ - [ ] Standardize embedding model paths (use full HuggingFace paths like "sentence-transformers/all-MiniLM-L6-v2" vs short names "all-MiniLM-L6-v2") - [ ] Clarify DBConfig constructor usage pattern (positional vs keyword arguments) -- [ ] Add cross-reference table showing C++ ↔ Python method equivalents -- [ ] Document GPUConfig differences between Python and C++ more prominently - [ ] Standardize index type string format (ivf_flat vs ivfflat) across JS/TS SDK ### API Documentation @@ -66,13 +59,11 @@ #### Service-specific - [ ] Add use cases to get-vector-count.mdx for consistency -- [ ] Document the extra Client constructor parameters in C++ (e.g., `0, false` parameters) ### General Improvements - [ ] Add "Common Mistakes" or "Common Pitfalls" sections to major methods - [ ] Verify JavaScript SDK package name `@cyborgdb/client` against actual npm package -- [ ] Verify C++ API against actual library implementation - [ ] Replace placeholder vector values `[1, 2, 3, ...]` with realistic examples or mark as pseudocode - [ ] Add more complex, real-world examples throughout - [ ] Verify embedding model short names vs. full paths support \ No newline at end of file diff --git a/versions/next/embedded/cpp/client/client.mdx b/versions/next/embedded/cpp/client/client.mdx deleted file mode 100644 index 82992e8..0000000 --- a/versions/next/embedded/cpp/client/client.mdx +++ /dev/null @@ -1,129 +0,0 @@ ---- -title: "Client" -mode: "wide" -noindex: true ---- - -The `cyborg::Client` class manages storage configurations and acts as a factory for creating or loading encrypted indexes. - -## Constructor - -```cpp -cyborg::Client(const std::string& api_key, - cyborg::StorageConfig storage_config, - const int cpu_threads, - const GPUConfig gpu_config); -``` -Initializes a new instance of `Client`. - -### Parameters - -| Parameter | Type | Description | -|-------------------|-----------------------|------------------------------------------------------| -| `api_key` | `std::string` | API key for your CyborgDB account. | -| `storage_config` | [`StorageConfig`](../types#storageconfig) | Backing store shared across the client and all of its indexes. Built via `StorageConfig::Disk(path)` or `StorageConfig::S3(bucket)`. | -| `cpu_threads` | `int` | Number of CPU threads to use (e.g., `0` to use all available cores).| -| `gpu_config` | [`GPUConfig`](../types#gpuconfig) | GPU operations configuration (requires CUDA). Use bitflags: `kNone` (no GPU), `kUpsert` (GPU for upsert), `kTrain` (GPU for training), `kQuery` (GPU for query), or `kAll` (GPU for all operations). Combine with `\|` operator.| - -`Client` is **non-copyable and non-movable** — it owns unique keystore handles. Construct it in place, or hold it via `std::unique_ptr`. - -### Exceptions - - - - - Throws if the `cpu_threads` parameter is less than `0`. - - Throws if the [`StorageConfig`](../types#storageconfig) is invalid. - - Throws if the GPU is not available when `gpu_accelerate` is `true`. - - - - Throws if the backing store is not available. - - Throws if the Client could not be initialized. - - - -### Example Usage - -```cpp -#include "cyborgdb_core/client.hpp" - -std::string api_key = "your_api_key_here"; // Replace with your CyborgDB API key -int cpu_threads = 4; - -// Local disk backing store (use StorageConfig::S3(bucket) for S3). -cyborg::StorageConfig storage = cyborg::StorageConfig::Disk("/tmp/cyborgdb"); - -// Example 1: Enable GPU for all operations -cyborg::Client client1(api_key, cyborg::StorageConfig::Disk("/tmp/cyborgdb"), cpu_threads, cyborg::kAll); - -// Example 2: Enable GPU for specific operations (upsert, train, and query) -cyborg::GPUConfig gpu_config_selective = cyborg::kUpsert | cyborg::kTrain | cyborg::kQuery; -cyborg::Client client2(api_key, cyborg::StorageConfig::Disk("/tmp/cyborgdb"), cpu_threads, gpu_config_selective); - -// Example 3: Disable GPU completely -cyborg::Client client3(api_key, cyborg::StorageConfig::Disk("/tmp/cyborgdb"), cpu_threads, cyborg::kNone); - -// Hold the client via unique_ptr (Client is non-movable) -auto client = std::make_unique( - api_key, cyborg::StorageConfig::Disk("/tmp/cyborgdb"), cpu_threads, cyborg::kNone); -``` - -Get your API key from the [API key page](../../../intro/get-api-key). - ---- - -## Methods - -### `cpu_threads()` - -```cpp -int cpu_threads() const; -``` - -Returns the number of CPU threads configured for this client. - -**Returns:** `int` - The number of CPU threads. - -**Example:** -```cpp -int threads = client.cpu_threads(); -std::cout << "Using " << threads << " CPU threads" << std::endl; -``` - ---- - -### `gpu_accelerate()` - -```cpp -bool gpu_accelerate() const; -``` - -Checks if GPU acceleration is enabled for any operations. - -**Returns:** `bool` - `true` if GPU is enabled for any operation, `false` otherwise. - -**Example:** -```cpp -if (client.gpu_accelerate()) { - std::cout << "GPU acceleration is enabled" << std::endl; -} -``` - ---- - -### `gpu_config()` - -```cpp -GPUConfig gpu_config() const; -``` - -Returns the GPU operations configuration for this client. - -**Returns:** [`GPUConfig`](../types#gpuconfig) - The GPU operations configuration. - -**Example:** -```cpp -cyborg::GPUConfig config = client.gpu_config(); -if (config & cyborg::kTrain) { - std::cout << "GPU is enabled for training" << std::endl; -} -``` \ No newline at end of file diff --git a/versions/next/embedded/cpp/client/create-index.mdx b/versions/next/embedded/cpp/client/create-index.mdx deleted file mode 100644 index 6149fef..0000000 --- a/versions/next/embedded/cpp/client/create-index.mdx +++ /dev/null @@ -1,96 +0,0 @@ ---- -title: "Create Index" -mode: "wide" -noindex: true ---- - -Creates and returns a new encrypted DiskIVF index. CyborgDB supports a single index type, DiskIVF; the older `IndexConfig`/`IndexIVF*` variants have been removed. - -```cpp -// Overload 1: Default configuration (DiskIVF, dimension auto-detected) -std::unique_ptr - CreateIndex(const std::string index_name, - const std::array& index_key, - const std::optional& metric = std::nullopt, - cyborg::Logger* logger = nullptr); - -// Overload 2: With an explicit IndexDiskIVF configuration -std::unique_ptr - CreateIndex(const std::string index_name, - const std::array& index_key, - const IndexDiskIVF& index_config, - const std::optional& metric = std::nullopt, - cyborg::Logger* logger = nullptr); -``` - -### Parameters - -| Parameter | Type | Description | -|----------------|-------------------------------|-----------------------------------------------------| -| `index_name` | `std::string` | Name of the index to create (must be unique). | -| `index_key` | `std::array` | 32-byte encryption key (KEK) for the index, used to secure index data. | -| `index_config` | [`IndexDiskIVF`](../types#indexdiskivf) | _(Overload 2 only)_ DiskIVF configuration (dimension, storage precision, embedding model). When omitted (overload 1), a default DiskIVF config is used and dimension is auto-detected from the first upsert. | -| `metric` | `std::optional` | _(Optional)_ Distance metric to override the one set in `index_config` (default is `std::nullopt`). | -| `logger` | [`cyborg::Logger*`](./logger) | _(Optional)_ Pointer to a logger instance for capturing operation logs (default is `nullptr`). | - -### Returns - -`std::unique_ptr`: A pointer to the newly created index ([`EncryptedIndex`](../encrypted-index)). - -### Exceptions - - - - - Throws if the index name is not unique. - - Throws if the index configuration is invalid. - - - - Throws if the index could not be created. - - - -### Example Usage - -#### Automatic Index Config - -```cpp -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include "cyborgdb_core/logger.hpp" -#include - -// ... Initialize the client ... - -// Create a secure 32-byte key (example: all zeros) -std::array index_key = {0}; - -auto index = client.CreateIndex("my_index", index_key); -``` - -#### Explicit Index Config - -```cpp -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include "cyborgdb_core/logger.hpp" -#include - -// ... Initialize the client ... - -// Create a secure 32-byte key (example: all zeros) -std::array index_key = {0}; - -// Example vector dimensionality -const size_t vector_dim = 1024; - -// Create a DiskIVF index configuration with an explicit dimension -cyborg::IndexDiskIVF index_config(vector_dim); - -// Optional: store rerank vectors in half precision to shrink the on-disk footprint -// cyborg::IndexDiskIVF index_config(vector_dim, "", cyborg::StoragePrecision::Float16); - -// Optional: Create and configure a logger -cyborg::Logger logger; -logger.Configure(LogLevel::Info, true, "index_creation.log"); - -auto index = client.CreateIndex("my_index", index_key, index_config, cyborg::DistanceMetric::Euclidean, &logger); \ No newline at end of file diff --git a/versions/next/embedded/cpp/client/list-indexes.mdx b/versions/next/embedded/cpp/client/list-indexes.mdx deleted file mode 100644 index 93f29e9..0000000 --- a/versions/next/embedded/cpp/client/list-indexes.mdx +++ /dev/null @@ -1,38 +0,0 @@ ---- -title: "List Indexes" -mode: "wide" -noindex: true ---- - -Returns a list of all encrypted index names accessible via the client at its configured [`StorageConfig`](../types#storageconfig). - -```cpp -std::vector ListIndexes(); -``` - -### Returns - -`std::vector`: A list of index names. - -### Exceptions - - - - - Throws if the list of indexes could not be retrieved. - - - -### Example Usage - -```cpp -#include "cyborgdb_core/client.hpp" - -// ... Initialize the client ... - -auto indexes = client.ListIndexes(); - -// Print the list of indexes -for (const auto& name : indexes) { - std::cout << name << std::endl; -} -``` \ No newline at end of file diff --git a/versions/next/embedded/cpp/client/load-index.mdx b/versions/next/embedded/cpp/client/load-index.mdx deleted file mode 100644 index 3c6ccac..0000000 --- a/versions/next/embedded/cpp/client/load-index.mdx +++ /dev/null @@ -1,58 +0,0 @@ ---- -title: "Load Index" -mode: "wide" -noindex: true ---- - -Loads an existing encrypted index and returns an instance of `EncryptedIndex`. - -```cpp -std::unique_ptr - LoadIndex(const std::string index_name, - const std::array& index_key, - cyborg::Logger* logger = nullptr, - std::optional> user_id = std::nullopt); -``` - -### Parameters - -| Parameter | Type | Description | -|----------------|-------------------------------|-----------------------------------------------------| -| `index_name` | `std::string` | Name of the index to load. | -| `index_key` | `std::array` | 32-byte encryption key (KEK) for the index. For a root load this is the root index KEK; for an RBAC user-scoped load (see `user_id`) this is that user's KEK. | -| `logger` | [`cyborg::Logger*`](./logger) | _(Optional)_ Pointer to a logger instance for capturing operation logs (default is `nullptr`). | -| `user_id` | `std::optional>` | _(Optional)_ 16-byte RBAC user identifier. When provided, the index is loaded user-scoped: `index_key` is treated as that user's KEK and per-operation permissions are enforced. Defaults to `std::nullopt` (root load). | - -### Returns - -`std::unique_ptr`: A pointer to the loaded index ([`EncryptedIndex`](../encrypted-index)). - -### Exceptions - - - - - Throws if the index name does not exist. - - - - Throws if the index could not be loaded or decrypted. - - - -### Example Usage - -```cpp -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include "cyborgdb_core/logger.hpp" -#include - -// ... Initialize the client ... - -// Create a secure 32-byte key (example: all zeros) -std::array index_key = {0}; - -// Optional: Create and configure a logger -cyborg::Logger logger; -logger.Configure(LogLevel::Debug, true, "index_loading.log"); - -auto index = client.LoadIndex("my_index", index_key, &logger); \ No newline at end of file diff --git a/versions/next/embedded/cpp/client/logger.mdx b/versions/next/embedded/cpp/client/logger.mdx deleted file mode 100644 index dce26ef..0000000 --- a/versions/next/embedded/cpp/client/logger.mdx +++ /dev/null @@ -1,44 +0,0 @@ ---- -title: "Logger" -mode: "wide" -noindex: true ---- - -The `cyborg::Logger` class provides thread-safe logging capabilities for CyborgDB C++ operations. It follows a singleton pattern, allowing you to configure and access the same logger instance across your application. - -## Configure Logger - -```cpp -void Configure(LogLevel level, - bool to_file, - const std::string& file_path = ""); -``` - -Configures the logger with the specified log level and output options. - -### Parameters - -| Parameter | Type | Default | Description | -|-----------|------|---------|-------------| -| `level` | `LogLevel` | - | Log level enum value: `LogLevel::Debug`, `LogLevel::Info`, `LogLevel::Warning`, `LogLevel::Error`, or `LogLevel::Critical`. | -| `to_file` | `bool` | - | Whether to write logs to a file instead of stdout/stderr. | -| `file_path` | `const std::string&` | `""` | _(Optional)_ Path to the log file. Only used when `to_file` is `true`. | - -### Exceptions - -- If the file cannot be opened, logs will be written to standard output instead, and an error message will be printed to `std::cerr`. - -### Example Usage - -```cpp -// Get the logger instance -cyborg::Logger& logger = cyborg::Logger::Instance(); - -// Configure logger to write INFO level logs to console -logger.Configure(LogLevel::Info, false); - -// Configure logger to write ERROR level logs to a file -logger.Configure(LogLevel::Error, true, "/path/to/logs/cyborgdb.log"); -``` - -The `Debug` level is only available in debug builds of the library (compiled with the DEBUG flag). \ No newline at end of file diff --git a/versions/next/embedded/cpp/encrypted-index/delete-index.mdx b/versions/next/embedded/cpp/encrypted-index/delete-index.mdx deleted file mode 100644 index c032425..0000000 --- a/versions/next/embedded/cpp/encrypted-index/delete-index.mdx +++ /dev/null @@ -1,38 +0,0 @@ ---- -title: "Delete Index" -mode: "wide" -noindex: true ---- - -This action is irreversible. Proceed with caution. - -Deletes the current index and all its associated data. - -```cpp -void DeleteIndex(const KeyContext& key); -``` - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `key` | [`KeyContext`](../types#keycontext) | Key context for the operation. Must be the **root** 32-byte index KEK. A bare 32-byte index key implicitly converts to a `KeyContext`. | - -`DeleteIndex` requires the root index KEK. A per-user `KeyContext` (one built with a `user_kek` and `user_id`) is rejected — RBAC users cannot delete the index. - -### Exceptions - - - - - Throws if the index could not be deleted. - - Throws if a per-user key context is supplied instead of the root KEK. - - - -### Example Usage - -```cpp -index->DeleteIndex(index_key); -``` - -`index_key` is the 32-byte `std::array` root index KEK and converts implicitly to a `KeyContext`. diff --git a/versions/next/embedded/cpp/encrypted-index/delete.mdx b/versions/next/embedded/cpp/encrypted-index/delete.mdx deleted file mode 100644 index e5bb1c6..0000000 --- a/versions/next/embedded/cpp/encrypted-index/delete.mdx +++ /dev/null @@ -1,37 +0,0 @@ ---- -title: "Delete" -mode: "wide" -noindex: true ---- - -This action is irreversible. Proceed with caution. - -Deletes the specified encrypted items stored in the index, including all its associated fields (`vector`, `contents`, `metadata`). - -```cpp -void Delete(const std::vector& ids, const KeyContext& key); -``` - -**Parameters**: -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `ids` | `const std::vector&` | IDs to delete. | -| `key` | [`KeyContext`](../types#keycontext) | Key context for the operation. A bare 32-byte index key (the `index_key`) implicitly converts to a `KeyContext`. For an RBAC user, pass `cyborg::KeyContext{user_kek, user_id}` (write permission required). | - -### Exceptions - - - - - Throws if the items could not be deleted. - - Throws if the supplied key context lacks write permission. - - - -### Example Usage - -```cpp -// Delete items with IDs "item_1" and "item_2" -index->Delete({"item_1", "item_2"}, index_key); -``` - -`index_key` is the 32-byte `std::array` index KEK and converts implicitly to a `KeyContext`. diff --git a/versions/next/embedded/cpp/encrypted-index/get.mdx b/versions/next/embedded/cpp/encrypted-index/get.mdx deleted file mode 100644 index 5ab404f..0000000 --- a/versions/next/embedded/cpp/encrypted-index/get.mdx +++ /dev/null @@ -1,51 +0,0 @@ ---- -title: "Get" -mode: "wide" -noindex: true ---- - -Retrieves and decrypts items associated with the specified IDs. - -```cpp -std::vector - Get(const std::vector& ids, - const std::vector& include, - const KeyContext& key); -``` - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `ids` | `const std::vector&` | IDs to retrieve. (For a single item, provide a `std::vector` with one element.)| -| `include` | [`std::vector`](../types#itemfields) | List of item fields to return. Can include `kVector`, `kContents`, and `kMetadata`. | -| `key` | [`KeyContext`](../types#keycontext) | Key context for the operation. A bare 32-byte index key (the `index_key`) implicitly converts to a `KeyContext`. For an RBAC user, pass `cyborg::KeyContext{user_kek, user_id}` (read permission required). | - -### Returns - -[`std::vector`](../types#item): Decrypted items with requested fields. - -IDs will always be included in the returned items. - -### Exceptions - - - - - Throws if the items could not be retrieved or decrypted. - - Throws if the supplied key context lacks read permission. - - - -### Example Usage - -```cpp -std::vector item_ids = {"item_1", "item_2"}; -auto items = index->Get(item_ids, {cyborg::ItemFields::kContents}, index_key); - -for (const auto& item : items) { - // Process each decrypted item (IDs & contents) - std::string id = item.id; -} -``` - -`index_key` is the 32-byte `std::array` index KEK and converts implicitly to a `KeyContext`. diff --git a/versions/next/embedded/cpp/encrypted-index/list-ids.mdx b/versions/next/embedded/cpp/encrypted-index/list-ids.mdx deleted file mode 100644 index 4342faa..0000000 --- a/versions/next/embedded/cpp/encrypted-index/list-ids.mdx +++ /dev/null @@ -1,57 +0,0 @@ ---- -title: "List IDs" -mode: "wide" -noindex: true ---- - -Lists all item IDs currently stored in the index. - -```cpp -std::vector ListIDs(const KeyContext& key); -``` - -The trailing `key` is the index's 32-byte KEK. A bare `std::array` implicitly converts to a `KeyContext`, so you can pass `index_key` directly; for an RBAC user pass `cyborg::KeyContext{user_kek, user_id}`. - -### Returns - -`std::vector`: A vector containing all item IDs currently stored in the encrypted index. - -### Exceptions - - - - - Throws if an error occurs during retrieval of IDs from the index. - - - -### Example Usage - -```cpp -#include "encrypted_index.hpp" - -// ... Create or load an encrypted index ... - -// Get all IDs in the index -auto all_ids = index->ListIDs(index_key); - -// Print all IDs -std::cout << "Index contains " << all_ids.size() << " items:" << std::endl; -for (const auto& id : all_ids) { - std::cout << "- " << id << std::endl; -} - -// Check if a specific ID exists -std::string target_id = "item_123"; -bool exists = std::find(all_ids.begin(), all_ids.end(), target_id) != all_ids.end(); -if (exists) { - std::cout << "Item " << target_id << " exists in the index" << std::endl; -} -``` - - -Use `ListIDs()` to: -- Verify what items are stored in your index -- Check if specific items exist before querying -- Get a complete inventory of your encrypted index contents -- Implement custom batch processing logic - \ No newline at end of file diff --git a/versions/next/embedded/cpp/encrypted-index/manage-users.mdx b/versions/next/embedded/cpp/encrypted-index/manage-users.mdx deleted file mode 100644 index 8cc5e1d..0000000 --- a/versions/next/embedded/cpp/encrypted-index/manage-users.mdx +++ /dev/null @@ -1,162 +0,0 @@ ---- -title: "Manage Users (RBAC)" -mode: "wide" -noindex: true ---- - -CyborgDB supports role-based access control (RBAC) through per-user keys. The holder of the -**root** index KEK can mint per-user wrapped keys that grant read and/or write access to an index. - -## How RBAC works - -- **Per-index KEK.** Each index is encrypted under a 32-byte Key-Encryption-Key (KEK), the - `index_key`. CyborgDB wraps a data key (DEK) under the KEK internally — callers only ever handle - the KEK. -- **Per-user keys.** The root KEK holder calls `CreateUserKeys` to wrap the index key under a - per-user 32-byte key (`user_kek`), granting read and/or write access. **The set of wraps that - exist defines the permission set** — there is no separate permission flag. A read-only user - simply has no write wrap. -- **Root-gated administration.** Creating, deleting, and listing user keys always requires the - root index KEK. -- **User load.** A user then opens the index with their own `user_kek`, scoped to their `user_id`: - ```cpp - auto index = client.LoadIndex(index_name, user_kek, /*logger=*/nullptr, user_id); - ``` - Per-operation, the gate is enforced — a read-only user's `Upsert`/`Delete` are rejected. - -Pass a per-user key context to data operations as `cyborg::KeyContext{user_kek, user_id}` -where `user_id` is a `std::array` and `user_kek` is a `std::array`. - ---- - -## CreateUserKeys - -```cpp -void CreateUserKeys(const std::array& user_id, - const std::array& index_key, - const std::array& user_kek, - bool grant_read, - bool grant_write); -``` - -Mints per-user wrapped keys granting the requested permissions. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `user_id` | `std::array` | 16-byte user identifier. | -| `index_key` | `std::array` | The **root** index KEK (admin gate). | -| `user_kek` | `std::array` | The user's 32-byte key under which access is wrapped. | -| `grant_read` | `bool` | Grant read access. | -| `grant_write` | `bool` | Grant write access. | - -At least one of `grant_read` / `grant_write` must be `true`. - -### Exceptions - - - - - Throws if the supplied key is not the root index KEK. - - Throws if the user keys could not be created. - - - ---- - -## DeleteUserKeys - -```cpp -void DeleteUserKeys(const std::array& user_id, - const std::array& index_key); -``` - -Revokes a user's access by deleting their wrapped keys. Idempotent and root-gated. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `user_id` | `std::array` | 16-byte user identifier to revoke. | -| `index_key` | `std::array` | The **root** index KEK (admin gate). | - -### Exceptions - - - - - Throws if the supplied key is not the root index KEK. - - - ---- - -## ListUserKeys - -```cpp -std::vector ListUserKeys(const std::array& index_key); -``` - -Lists all users that have keys for the index and their permissions. Root-gated. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `index_key` | `std::array` | The **root** index KEK (admin gate). | - -### Returns - -`std::vector`: One entry per user. The `UserKeyInfo` struct: - -```cpp -struct UserKeyInfo { - std::array user_id; - bool has_read; - bool has_write; -}; -``` - -### Exceptions - - - - - Throws if the supplied key is not the root index KEK. - - - ---- - -## Example Usage - -```cpp -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include -#include - -// ... Initialize the client and load/create the index with the root index_key ... - -// 16-byte user identifier and the user's 32-byte key (16/32 random bytes each) -std::array user_id; -std::array user_kek; -RAND_bytes(user_id.data(), user_id.size()); -RAND_bytes(user_kek.data(), user_kek.size()); - -// Admin (root KEK holder) grants the user read-only access -index->CreateUserKeys(user_id, index_key, user_kek, /*grant_read=*/true, /*grant_write=*/false); - -// Inspect current users -for (const auto& info : index->ListUserKeys(index_key)) { - std::cout << "read=" << info.has_read << " write=" << info.has_write << "\n"; -} - -// The user loads the index scoped to their identity -auto user_index = client.LoadIndex("my_index", user_kek, /*logger=*/nullptr, user_id); - -// Per-operation, the user passes a per-user key context -cyborg::Array2D q{{0.1f, 0.2f, 0.3f, 0.4f}}; -auto results = user_index->Query(q, cyborg::QueryParams{}, cyborg::KeyContext{user_kek, user_id}); - -// Revoke access when done -index->DeleteUserKeys(user_id, index_key); -``` diff --git a/versions/next/embedded/cpp/encrypted-index/query.mdx b/versions/next/embedded/cpp/encrypted-index/query.mdx deleted file mode 100644 index bd04d45..0000000 --- a/versions/next/embedded/cpp/encrypted-index/query.mdx +++ /dev/null @@ -1,63 +0,0 @@ ---- -title: "Query" -mode: "wide" -noindex: true ---- - -Retrieves the nearest neighbors for given query vectors. - -```cpp -QueryResults Query(Array2D& query_vectors, - const QueryParams& query_params, - const KeyContext& key); -``` - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `query_vectors` | [`Array2D`](../types#array2d) | Query vectors to search. | -| `query_params` | [`QueryParams`](../types#queryparams) | Parameters for querying, such as `top_k`, `n_probes`, and `rerank_mult` (default 50). Pass `cyborg::QueryParams{}` for defaults. | -| `key` | [`KeyContext`](../types#keycontext) | Key context for the operation. A bare 32-byte index key (the `index_key`) implicitly converts to a `KeyContext`. For an RBAC user, pass `cyborg::KeyContext{user_kek, user_id}` (read permission required). | - -Both `query_params` and `key` are required. There are no longer default values for these arguments. - -If this function is called on an index where `TrainIndex()` has not been executed (or while a retrain is in progress), the query will use encrypted exhaustive search. -This may cause queries to be slower, especially when there are many vector embeddings in the index. - -### Returns - -[`QueryResults`](../types#queryresults): Results `top_k` containing decrypted nearest neighbors IDs and distances. - -### Exceptions - - - - - Throws if the query vectors have incompatible dimensions with the index. - - Throws if the index was not created or loaded yet. - - - - Throws if the query could not be executed. - - Throws if the supplied key context lacks read permission. - - - -### Example Usage - -```cpp -cyborg::Array2D queries{/*...*/}; // Populate with one or more query vectors - -// Query with default parameters -QueryResults results = index->Query(queries, cyborg::QueryParams{}, index_key); - -// Query with explicit parameters: top_k = 10, n_probes = 5 -cyborg::QueryParams params(10, 5); -QueryResults results2 = index->Query(queries, params, index_key); - -auto view = results[0]; // Access first query's results -for (uint32_t j = 0; j < view.num_results; ++j) { - std::cout << "ID: " << view.ids[j] << ", Distance: " << view.distances[j] << std::endl; -} -``` - -`index_key` is the 32-byte `std::array` index KEK and converts implicitly to a `KeyContext`. In a stateless service that reloads the index per request, pass the key on every operation. diff --git a/versions/next/embedded/cpp/encrypted-index/train.mdx b/versions/next/embedded/cpp/encrypted-index/train.mdx deleted file mode 100644 index efa78a1..0000000 --- a/versions/next/embedded/cpp/encrypted-index/train.mdx +++ /dev/null @@ -1,42 +0,0 @@ ---- -title: "Train Index" -mode: "wide" -noindex: true ---- - -Builds the index using the specified training configuration. Required before efficient querying. -Prior to calling this, all queries will be conducted using encrypted exhaustive search. -After, they will be conducted using encrypted ANN search. - -```cpp -void TrainIndex(const TrainingConfig& training_config, const KeyContext& key); -``` - -### Parameters - -| Parameter | Type | Description | -|----------|------------|--------------------| -| `training_config` | [`TrainingConfig`](../types#trainingconfig) | Training parameters including n_lists, batch size, max iterations, tolerance, and max memory. | -| `key` | [`KeyContext`](../types#keycontext) | Key context for the operation. A bare 32-byte index key (the `index_key`) implicitly converts to a `KeyContext`. | - -There must be at least `2 * n_lists` or `10,000` (whichever is greater) vector embeddings in the index prior to calling this function. The `n_lists` parameter is configured within the `TrainingConfig` object. - -### Exceptions - - - - - Throws if there are not enough vector embeddings in the index for training (must be at least `2 * n_lists`). - - Throws if the index could not be trained. - - - -### Example Usage - -```cpp -// TrainingConfig(n_lists, batch_size, max_iters, tolerance, max_memory) -// batch_size of 0 (auto) lets CyborgDB choose the batch size. -cyborg::TrainingConfig config(1024, 0, 100, 1e-6, 0); -index->TrainIndex(config, index_key); -``` - -`index_key` is the 32-byte `std::array` index KEK and converts implicitly to a `KeyContext`. diff --git a/versions/next/embedded/cpp/encrypted-index/upsert.mdx b/versions/next/embedded/cpp/encrypted-index/upsert.mdx deleted file mode 100644 index fbb2cc4..0000000 --- a/versions/next/embedded/cpp/encrypted-index/upsert.mdx +++ /dev/null @@ -1,65 +0,0 @@ ---- -title: "Upsert" -mode: "wide" -noindex: true ---- - -Adds or updates vector embeddings in the index. If an item already exists at `id`, then it will be overwritten. - -```cpp -void Upsert(const std::vector& ids, - Array2D& vectors, - const std::vector>& contents, - const std::vector& json_metadata_array, - const KeyContext& key); -``` - -### Parameters - -| Parameter | Type | Description | -|---------------|-----------------------------|-----------------------------------------------------| -| `ids` | `std::vector&` | Unique identifiers for each vector. | -| `vectors` | [`Array2D`](../types#array2d) | 2D container with vector embeddings to index. | -| `contents` | `std::vector>&` | Item contents in bytes. Pass `{}` if none. | -| `json_metadata_array` | `std::vector&` | Item metadata as serialized JSON strings. Pass `{}` if none. | -| `key` | [`KeyContext`](../types#keycontext) | Key context for the operation. A bare 32-byte index key (the `index_key`) implicitly converts to a `KeyContext`. For an RBAC user, pass `cyborg::KeyContext{user_kek, user_id}` (write permission required). | - -For more info on metadata, see [Metadata Filtering](../../guides/data-operations/metadata-filtering). - -`contents` and `json_metadata_array` are required positional arguments. Pass an empty initializer `{}` when an item has no contents or metadata. When provided, their length must match `ids`. - -### Exceptions - - - - - Throws if vector dimensions are incompatible with the index configuration. - - Throws if index was not created or loaded yet. - - Throws if there is a mismatch between the number of `vectors`, `ids`, `contents` or `json_metadata_array`. - - - - Throws if the vectors could not be upserted. - - Throws if the supplied key context lacks write permission. - - - -### Example Usage - -```cpp -cyborg::Array2D embeddings{{0.1, 0.2, 0.3}, {0.4, 0.5, 0.6}}; -std::vector ids = {"item_1", "item_2"}; - -// Upsert without contents or metadata (pass {} for both) -index->Upsert(ids, embeddings, {}, {}, index_key); - -// Upsert with associated item contents -std::vector> contents = { - {'a', 'b', 'c'}, {'d', 'e', 'f'} -}; -index->Upsert(ids, embeddings, contents, {}, index_key); - -// Upsert with contents and metadata -std::vector metadata = {"{\"type\": \"image\"}", "{\"type\": \"text\"}"}; -index->Upsert(ids, embeddings, contents, metadata, index_key); -``` - -`index_key` is the 32-byte `std::array` index KEK. It converts implicitly to a `KeyContext`. In a stateless service that reloads the index per request, pass the key on every operation. diff --git a/versions/next/embedded/cpp/getters.mdx b/versions/next/embedded/cpp/getters.mdx deleted file mode 100644 index 49d69f0..0000000 --- a/versions/next/embedded/cpp/getters.mdx +++ /dev/null @@ -1,131 +0,0 @@ ---- -title: "Getter Functions" -noindex: true ---- - -## is_trained - -```cpp -bool is_trained() const; -``` - -Returns whether the index has been trained. Returns `true` only when the training state is -`Trained`. While a (re)train rebuild is in progress this returns `false`, and queries -transparently fall back to the untrained (exhaustive) search path. See -[`is_training`](#is_training) and [`training_state`](#training_state). - ---- - -## is_training - -```cpp -bool is_training() const; -``` - -Returns `true` while a (re)train rebuild is in progress, otherwise `false`. - ---- - -## training_state - -```cpp -TrainingState training_state() const; -``` - -Returns the current training state as a [`TrainingState`](./types#trainingstate) enum: -`Untrained`, `Training`, or `Trained`. - ---- - -## index_name - -```cpp -std::string index_name() const; -``` - -Returns the name of the index. - ---- - -## index_type - -```cpp -IndexType index_type() const; -``` - -Returns the type of the index as an [`IndexType`](./types#indextype) enum. The only value is -`DISK_IVF`. - ---- - -## index_config - -```cpp -IndexDiskIVF* index_config() const; -``` - -Returns a pointer to the index configuration ([`IndexDiskIVF`](./types#indexdiskivf)). - ---- - -## ListIDs - -```cpp -std::vector ListIDs(const KeyContext& key); -``` - -Returns a vector containing all item IDs currently stored in the index. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `key` | [`KeyContext`](./types#keycontext) | Key context for the operation. A bare 32-byte index key (the `index_key`) implicitly converts to a `KeyContext`. | - -### Exceptions - -- **std::runtime_error**: Thrown if an error occurs during retrieval. - ---- - -## NumVectors - -```cpp -size_t NumVectors(const KeyContext& key); -``` - -Returns the total number of vectors currently stored in the encrypted index. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `key` | [`KeyContext`](./types#keycontext) | Key context for the operation. A bare 32-byte index key (the `index_key`) implicitly converts to a `KeyContext`. | - -### Returns - -`size_t`: The number of vectors in the index. - -### Exceptions - - - - - Throws if the index was not created or loaded yet. - - Throws if an error occurs while retrieving the count. - - - -### Example Usage - -```cpp -// Get the number of vectors in the index -size_t vector_count = index->NumVectors(index_key); -std::cout << "Index contains " << vector_count << " vectors" << std::endl; - -// Use count for validation or progress reporting -if (vector_count == 0) { - std::cout << "Index is empty, consider adding vectors" << std::endl; -} -``` - -`index_key` is the 32-byte `std::array` index KEK and converts implicitly to a `KeyContext`. diff --git a/versions/next/embedded/cpp/introduction.mdx b/versions/next/embedded/cpp/introduction.mdx deleted file mode 100644 index 803fbda..0000000 --- a/versions/next/embedded/cpp/introduction.mdx +++ /dev/null @@ -1,14 +0,0 @@ ---- -title: "CyborgDB Embedded - C++ API Reference" -sidebarTitle: "C++ API Introduction" -description: "`v0.17.x`" -mode: "wide" -noindex: true ---- - -The C++ API for CyborgDB is split into two main classes within the `cyborg` namespace: - -- `Client` – Handles configuration, index creation/loading, and listing available indexes. -- `EncryptedIndex` – Provides data operations on a specific encrypted index such as upserting vectors, training the index, querying, and retrieving stored items. - -This API is also exposed via PyBind11 in the [Python API](../python/introduction). diff --git a/versions/next/embedded/cpp/kms.mdx b/versions/next/embedded/cpp/kms.mdx deleted file mode 100644 index 95afb88..0000000 --- a/versions/next/embedded/cpp/kms.mdx +++ /dev/null @@ -1,158 +0,0 @@ ---- -title: "KMS Key Management" -mode: "wide" -noindex: true ---- - -These module-level functions persist a per-index **KMS envelope** describing how the index KEK is -wrapped by an external Key Management Service. The envelope is stored as a FlatBuffer next to the -index keystore. - -KMS key wrapping is primarily for the **service layer**. Embedded SDK users who supply their -own KEK directly can ignore these functions, or use `provider = "none"`. They are advanced and -optional. - -```cpp -#include "cyborgdb_core/index_kms.hpp" -``` - -## Concepts - -- **KMS envelope.** A [`KMSBlob`](./types#kmsblob) records how an index's KEK is wrapped by an - external KMS. Callers never store the plaintext KEK in the envelope; they store the wrapped form - plus the metadata needed to unwrap it. -- **Providers.** - - `"aws"` — the KEK is wrapped with AES-256-GCM under a value stored in AWS Secrets Manager. - - `"aws-kms"` — the KEK is wrapped via `kms.Encrypt`. - - `"none"` — nothing is stored; the SDK supplies the plaintext KEK per request. -- **Strict insert vs. upsert.** `CreateIndexKMS` fails if an envelope already exists for the index; - `PushIndexKMS` is an idempotent upsert. - ---- - -## CreateIndexKMS - -```cpp -void CreateIndexKMS(const StorageConfig& config_location, - const std::string& index_name, - const KMSBlob& blob); -``` - -Stores a new KMS envelope for the index. Strict insert — throws if one already exists. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `config_location` | [`StorageConfig`](./types#storageconfig) | Backing store where the envelope is persisted. | -| `index_name` | `const std::string&` | Name of the index. | -| `blob` | [`KMSBlob`](./types#kmsblob) | The KMS envelope to store. | - ---- - -## PushIndexKMS - -```cpp -void PushIndexKMS(const StorageConfig& config_location, - const std::string& index_name, - const KMSBlob& blob); -``` - -Idempotent upsert of the KMS envelope for the index. Creates it if absent, overwrites it otherwise. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `config_location` | [`StorageConfig`](./types#storageconfig) | Backing store where the envelope is persisted. | -| `index_name` | `const std::string&` | Name of the index. | -| `blob` | [`KMSBlob`](./types#kmsblob) | The KMS envelope to store. | - ---- - -## GetIndexKMS - -```cpp -KMSBlob GetIndexKMS(const StorageConfig& config_location, - const std::string& index_name); -``` - -Retrieves the stored KMS envelope for the index. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `config_location` | [`StorageConfig`](./types#storageconfig) | Backing store where the envelope lives. | -| `index_name` | `const std::string&` | Name of the index. | - -### Returns - -[`KMSBlob`](./types#kmsblob): The stored KMS envelope. - ---- - -## DeleteIndexKMS - -```cpp -void DeleteIndexKMS(const StorageConfig& config_location, - const std::string& index_name); -``` - -Deletes the stored KMS envelope for the index. Idempotent. - -### Parameters - -| Parameter | Type | Description | -|-------------------|--------------------|-----------------------| -| `config_location` | [`StorageConfig`](./types#storageconfig) | Backing store where the envelope lives. | -| `index_name` | `const std::string&` | Name of the index. | - ---- - -## KMSBlob - -The KMS envelope struct (see [`KMSBlob`](./types#kmsblob)): - -```cpp -struct KMSBlob { - std::string kms_name; - std::string provider; // "aws" | "aws-kms" | "none" - std::string key_id; - std::string region; - std::vector wrapped_kek; - uint32_t version = 0; - int64_t created_at = 0; // unix epoch seconds -}; -``` - ---- - -## Example Usage - -```cpp -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/index_kms.hpp" - -auto config_location = cyborg::StorageConfig::Disk("/tmp/cyborgdb"); - -cyborg::KMSBlob blob; -blob.kms_name = "my-kms"; -blob.provider = "aws-kms"; -blob.key_id = "arn:aws:kms:us-east-1:123456789012:key/abcd"; -blob.region = "us-east-1"; -blob.wrapped_kek = {/* wrapped KEK bytes */}; -blob.version = 1; - -// Strict insert -cyborg::CreateIndexKMS(config_location, "my_index", blob); - -// Idempotent upsert -cyborg::PushIndexKMS(config_location, "my_index", blob); - -// Retrieve -cyborg::KMSBlob stored = cyborg::GetIndexKMS(config_location, "my_index"); - -// Delete -cyborg::DeleteIndexKMS(config_location, "my_index"); -``` diff --git a/versions/next/embedded/cpp/types.mdx b/versions/next/embedded/cpp/types.mdx deleted file mode 100644 index 861fb59..0000000 --- a/versions/next/embedded/cpp/types.mdx +++ /dev/null @@ -1,561 +0,0 @@ ---- -title: "Types" -noindex: true ---- - -## StorageConfig - -`StorageConfig` defines the backing store for an index and all of its per-index keystores. It is immutable and has no public default constructor — instances are created via the static factory methods below. A single `StorageConfig` is shared across the client and all indexes it manages. - -CyborgDB supports two backing stores: local disk and S3 (or any S3-compatible store such as MinIO). - -### Static Factories - -```cpp -static StorageConfig Disk(std::optional path, - CachePolicy cache_policy = {}); -static StorageConfig S3(std::string bucket, S3Options opts = {}); -``` - -| Factory | Description | -|---------|-------------| -| `Disk(path, cache_policy)` | Local persistent storage on disk at `path`. Pass a [`CachePolicy`](#cachepolicy) to keep hot data in memory. | -| `S3(bucket, opts)` | AWS S3 or any S3-compatible object store. Configure region, endpoint, prefix, and credentials via [`S3Options`](#s3options). | - -### Example Usage - -```cpp -#include "cyborgdb_core/client.hpp" - -// Local disk store with vector caching -cyborg::CachePolicy cache; -cache.vectors = true; -cyborg::StorageConfig disk = cyborg::StorageConfig::Disk("/tmp/cyborgdb", cache); - -// S3 store with explicit credentials -cyborg::S3Options opts; -opts.region = "us-east-1"; -opts.credentials = cyborg::S3Credentials{"ACCESS_KEY", "SECRET_KEY"}; -cyborg::StorageConfig s3 = cyborg::StorageConfig::S3("my-bucket", opts); -``` - -For more info, you can read about supported backing stores [here](../../intro/backing-stores). - ---- - -## CachePolicy - -`CachePolicy` controls which categories of data a disk-backed store keeps cached in memory for faster access. - -```cpp -struct CachePolicy { - bool vectors = false; // Cache vector data in memory - bool metadata = false; // Cache metadata in memory - bool ids = false; // Cache item IDs in memory -}; -``` - ---- - -## S3Options - -`S3Options` configures an S3 backing store created via [`StorageConfig::S3`](#storageconfig). - -```cpp -struct S3Options { - std::string prefix = ""; // Key prefix within the bucket - std::optional region; // AWS region - std::optional endpoint; // Custom endpoint (MinIO/Ceph/R2) - std::optional credentials; // Explicit credentials -}; -``` - -| Field | Type | Description | -|-------|------|-------------| -| `prefix` | `std::string` | _(Optional)_ Key prefix applied to all objects within the bucket. Defaults to `""`. | -| `region` | `std::optional` | _(Optional)_ AWS region. | -| `endpoint` | `std::optional` | _(Optional)_ Custom S3 endpoint for S3-compatible stores (MinIO, Ceph, Cloudflare R2). | -| `credentials` | `std::optional` | _(Optional)_ Explicit S3 credentials. Omit to use the AWS default credential provider chain (environment variables, `~/.aws/credentials`, EC2 instance profile, EKS IRSA). | - -Path-style addressing is selected automatically when `endpoint` is set (MinIO/Ceph/R2); otherwise virtual-hosted addressing is used. - ---- - -## S3Credentials - -`S3Credentials` holds explicit credentials for an S3 backing store. - -```cpp -struct S3Credentials { - std::string access_key; - std::string secret_key; - std::optional session_token; -}; -``` - -| Field | Type | Description | -|-------|------|-------------| -| `access_key` | `std::string` | AWS access key ID. | -| `secret_key` | `std::string` | AWS secret access key. | -| `session_token` | `std::optional` | _(Optional)_ Session token for temporary credentials. | - ---- - -## GPUConfig - -`GPUConfig` is an enum that specifies which operations should use GPU acceleration. It uses bitflags that can be combined using the `|` (OR) operator. - -### Enum Values - -```cpp -enum GPUConfig : uint8_t { - kNone = 0, // No GPU usage - kUpsert = 1 << 0, // Use GPU for upsert operations - kTrain = 1 << 1, // Use GPU for training operations - kQuery = 1 << 2, // Use GPU for query operations - kAll = kUpsert | kTrain | kQuery // Use GPU for all operations -}; -``` - -### Example Usage - -```cpp -// Enable GPU for all operations -cyborg::GPUConfig config1 = cyborg::kAll; - -// Enable GPU only for training and query -cyborg::GPUConfig config2 = cyborg::kTrain | cyborg::kQuery; - -// Enable GPU only for upsert -cyborg::GPUConfig config3 = cyborg::kUpsert; - -// Disable GPU completely -cyborg::GPUConfig config4 = cyborg::kNone; -``` - ---- - -## DeviceConfig - -`DeviceConfig` class holds the configuration details for the device used in vector search operations, such as the number of CPU threads and GPU acceleration settings. - -### Constructor - -```cpp -DeviceConfig(const int cpu_threads = 0, const GPUConfig gpu_config = kNone); -``` - -### Parameters - -| Parameter | Type | Description | -|------------------------|-------------------------|-------------------------| -| `cpu_threads` | `int` | _(Optional)_ Number of CPU threads to use. Defaults to `0` (use all available cores). | -| `gpu_config` | [`GPUConfig`](#gpuconfig) | _(Optional)_ GPU operations configuration. Defaults to `kNone` (no GPU). | - -### Methods - -| Method | Return Type | Description | -|------------------------|-------------------------|-------------------------| -| `cpu_threads() const` | `int` | Get the number of CPU threads configured. | -| `gpu_config() const` | [`GPUConfig`](#gpuconfig) | Get the GPU operations configuration. | - -### Example Usage - -```cpp -// 4 CPU threads, GPU enabled for training and query -cyborg::DeviceConfig device_config(4, cyborg::kTrain | cyborg::kQuery); -int threads = device_config.cpu_threads(); // Returns 4 -cyborg::GPUConfig gpu = device_config.gpu_config(); // Returns kTrain | kQuery -``` - ---- - -## DistanceMetric - -The `DistanceMetric` enum contains the supported distance metrics for CyborgDB. These are: - -```cpp -enum class DistanceMetric { - Cosine, - Euclidean, - SquaredEuclidean}; -``` - ---- - -## IndexDiskIVF - -`IndexDiskIVF` configures a DiskIVF index — the single index type supported in CyborgDB. It replaces the older `IndexConfig` family. Pass an instance to [`CreateIndex`](./client/create-index) when you want explicit control over dimensionality or storage precision; otherwise the default-config overload of `CreateIndex` constructs one for you. - -### Constructor - -```cpp -IndexDiskIVF(size_t dimension = 0, - std::optional embedding_model = "", - StoragePrecision storage_precision = StoragePrecision::Float32); -``` - -### Parameters - -| Parameter | Type | Default | Description | -|--------------|-------------------------|---------|---------------------------------------| -| `dimension` | `size_t` | `0` | _(Optional)_ Dimensionality of vector embeddings. Auto-detected from the first upsert if `0`. | -| `embedding_model` | `std::optional` | `""` | _(Optional)_ Embedding model name for auto-generation; dimension can be derived from it. | -| `storage_precision` | [`StoragePrecision`](#storageprecision) | `Float32` | _(Optional)_ On-disk dtype of rerank vectors. `Float16` halves the disk footprint with a slight precision loss. | - -### Methods - -| Method | Return Type | Description | -|--------------|-------------------------|---------------------------------------| -| `dimension()` | `size_t` | Get vector dimensionality. | -| `set_dimension(size_t)` | `void` | Set vector dimensionality. | -| `metric()` | [`DistanceMetric`](#distancemetric) | Get distance metric. | -| `set_metric(DistanceMetric)` | `void` | Set distance metric. | -| `index_type()` | [`IndexType`](#indextype) | Returns `DISK_IVF`. | -| `embedding_model()` | `std::optional` | Get the embedding model name. | -| `storage_precision()` | [`StoragePrecision`](#storageprecision) | Get the on-disk storage precision. | -| `set_storage_precision(StoragePrecision)` | `void` | Set the on-disk storage precision. | -| `n_lists()` | `size_t` | Get number of inverted lists (initially 1, set during training). | -| `set_n_lists(size_t)` | `void` | Set number of inverted lists (usually done automatically during training). | - -### Example Usage - -```cpp -// Default configuration (dimension auto-detected, float32 rerank vectors) -cyborg::IndexDiskIVF config1; - -// Explicit dimension -cyborg::IndexDiskIVF config2(1024); - -// Explicit dimension with float16 storage precision (smaller on-disk footprint) -cyborg::IndexDiskIVF config3(1024, "", cyborg::StoragePrecision::Float16); -``` - ---- - -## StoragePrecision - -`StoragePrecision` controls the on-disk dtype of rerank vectors for a DiskIVF index. - -```cpp -enum class StoragePrecision { - Float32, // Full precision (default) - Float16 // Half precision — halves disk footprint, slight precision loss -}; -``` - ---- - -## TrainingState - -`TrainingState` reports the lifecycle state of an index's training. - -```cpp -enum class TrainingState : uint8_t { - Untrained = 0, // Index has not been trained - Training = 1, // A (re)train rebuild is in progress - Trained = 2 // Training is complete -}; -``` - -While an index is in the `Training` state, queries transparently fall back to the untrained (exhaustive) path. - ---- - -## IndexType - -The `IndexType` enum defines the supported index types in CyborgDB. CyborgDB now supports a single index type, DiskIVF: - -```cpp -enum IndexType { - DISK_IVF -}; -``` - ---- - -## Array2D - -`Array2D` class provides a 2D container for data, which can be initialized with a specific number of rows and columns, or from an existing vector. - -### Constructors - -```cpp -Array2D(size_t rows, size_t cols, const T& initial_value = T()); -Array2D(std::vector&& data, size_t cols); -Array2D(const std::vector& data, size_t cols); -Array2D(std::initializer_list> init_list); -Array2D(Array2D&& other) noexcept; -Array2D(); -``` - -- **`Array2D(size_t rows, size_t cols, const T& initial_value = T())`**: Creates a 2D array with specified dimensions, initialized with the given value. -- **`Array2D(std::vector&& data, size_t cols)`**: Initializes the 2D array from a 1D vector (move semantics). -- **`Array2D(const std::vector& data, size_t cols)`**: Initializes the 2D array from a 1D vector (copy). -- **`Array2D(std::initializer_list> init_list)`**: Initializes from a nested initializer list (e.g., `{{1, 2}, {3, 4}}`). -- **`Array2D(Array2D&& other) noexcept`**: Move constructor - transfers ownership without copying. -- **`Array2D()`**: Default constructor - creates an empty array (0 rows, 0 columns). - -The copy constructor is deleted. Use `Clone()` or move semantics to copy an `Array2D`. - -### Access Methods - -- **`operator()(size_t row, size_t col) const`**: Access an element at the specified row and column (read-only). -- **`operator()(size_t row, size_t col)`**: Access an element at the specified row and column (read-write). -- **`size_t rows() const`**: Returns the number of rows. -- **`size_t cols() const`**: Returns the number of columns. -- **`size_t size() const`**: Returns the total number of elements. - -### Example Usage - -```cpp -// Converting a vector to an array -std::vector vec = {0, 1, 2, 3, 4, 5, 6, 7}; -cyborg::Array2D arr(vec, 2); -// arr is now a 2D array of 4 rows and 2 columns, with the contents from vec - -// Creating a 2D array with 3 rows and 2 columns, initialized to zero -cyborg::Array2D array(3, 2, 0); - -// Access and modify elements -array(0, 0) = 1; -array(0, 1) = 2; - -// Printing the array -for (size_t i = 0; i < array.rows(); ++i) { - for (size_t j = 0; j < array.cols(); ++j) { - std::cout << array(i, j) << " "; - } - std::cout << std::endl; -} -``` - ---- - -## TrainingConfig - -The `TrainingConfig` struct defines parameters for training an index, allowing control over convergence and memory usage. - -### Constructor - -```cpp -TrainingConfig(std::optional n_lists = std::nullopt, - std::optional batch_size = std::nullopt, - std::optional max_iters = std::nullopt, - std::optional tolerance = std::nullopt, - std::optional max_memory = std::nullopt); -``` - -### Parameters - -| Parameter | Type | Description | -|-------------------|----------|---------------------------------------------------------------------------------------------| -| `n_lists` | `std::optional` | _(Optional)_ Number of inverted lists to create. Defaults to `std::nullopt` (auto-determines, typically `0`). | -| `batch_size` | `std::optional` | _(Optional)_ Size of each batch for training. Defaults to `std::nullopt` (auto-determined based on dataset size). | -| `max_iters` | `std::optional` | _(Optional)_ Maximum iterations for training. Defaults to `std::nullopt` (auto-determines, typically `100`). | -| `tolerance` | `std::optional` | _(Optional)_ Convergence tolerance for training. Defaults to `std::nullopt` (uses `1e-6`). | -| `max_memory` | `std::optional` | _(Optional)_ Maximum memory (MB) usage during training. Defaults to `std::nullopt` (no limit). | - -### Struct Members - -Note: The struct members are stored in this order (different from constructor parameter order): - -```cpp -size_t batch_size; // Batch size (default: 0, auto) -size_t max_iters; // Maximum iterations (default: 100) -double tolerance; // Convergence tolerance (default: 1e-6) -size_t max_memory; // Maximum memory in MB (default: 0, no limit) -size_t n_lists; // Number of inverted lists (default: 0, auto-determine) -``` - ---- - -## QueryParams - -The `QueryParams` struct defines parameters for querying the index, controlling the number of results, probing behavior, and reranking. - -### Constructor - -```cpp -explicit QueryParams(size_t top_k = 100, - size_t n_probes = 0, - std::string filters = "", - std::vector include = {}, - bool greedy = false, - size_t rerank_mult = 50); // kDefaultRerankMult = 50 -``` - -### Parameters - -| Parameter | Type | Description | -|-------------------|----------|--------------------------------------------------------------------------| -| `top_k` | `size_t` | _(Optional)_ Number of nearest neighbors to return. Defaults to `100`. | -| `n_probes` | `size_t` | _(Optional)_ Number of lists to probe during query. Defaults to `0` which will auto-determine optimal probes. | -| `filters` | `std::string` | _(Optional)_ A JSON string of filters to apply to vector metadata, limiting search scope to these vectors. | -| `include` | `std::vector` | _(Optional)_ List of result fields to return. Can include `kDistance` and `kMetadata`. Defaults to empty. | -| `greedy` | `bool` | _(Optional)_ Whether to perform greedy search. Defaults to `false`. | -| `rerank_mult` | `size_t` | _(Optional)_ Stage-1 retrieval multiplier used for reranking on DiskIVF indexes. Defaults to `50` (`kDefaultRerankMult`). | - -Higher n_probes values may improve recall but could slow down query time, so select a value based on desired recall and performance trade-offs. - -`filters` use a subset of the [MongoDB Query and Projection Operators](https://www.mongodb.com/docs/manual/reference/operator/query/). -For instance: `filters: { "$and": [ { "label": "cat" }, { "confidence": { "$gte": 0.9 } } ] }` means that only vectors where `label == "cat"` and `confidence >= 0.9` will be considered for encrypted vector search. -For more info on metadata, see [Metadata Filtering](../guides/data-operations/metadata-filtering). - ---- - -### QueryResults - -`QueryResults` class holds the results from a `Query` operation, including IDs, distances, and metadata for the nearest neighbors of each query. Results are vector-based and immutable after construction. - -### Getter Methods - -| Method | Return Type | Description | -|--------------------------|------------------------------|------------------------------------------------------| -| `ids()` | `const std::vector>&` | IDs of nearest neighbors for each query. | -| `distances()` | `const std::vector>&` | Distances of nearest neighbors for each query. | -| `metadata()` | `const std::vector>&` | Metadata for nearest neighbors for each query (JSON strings). | - -### Methods - -| Method | Return Type | Description | -|--------------------------|------------------------------|------------------------------------------------------| -| `ResultView operator[](size_t query_idx) const` | `ResultView` | Returns a read-only view of IDs, distances, and metadata for a specific query. | -| `num_results() const` | `std::vector` | Returns the actual number of results per query (may be less than top_k). | -| `num_queries() const` | `size_t` | Returns the number of queries. | -| `bool empty() const` | `bool` | Checks if the results are empty. | -| `static QueryResults Empty(size_t num_queries)` | `QueryResults` | Factory method to create empty results for a given number of queries. | - -### ResultView - -The `ResultView` struct provides read-only access to results for a single query: - -```cpp -struct ResultView { - const std::vector& ids; - const std::vector& distances; - const std::vector& metadata; - const uint32_t& num_results; -}; -``` - -### Example Usage - -```cpp -// Access results for each query -for (size_t i = 0; i < results.num_queries(); ++i) { - auto view = results[i]; - for (uint32_t j = 0; j < view.num_results; ++j) { - std::cout << "ID: " << view.ids[j] - << ", Distance: " << view.distances[j] << std::endl; - } -} - -// Access all IDs and distances directly -const auto& all_ids = results.ids(); -const auto& all_distances = results.distances(); - -// Get actual result counts per query -auto counts = results.num_results(); - -// Create empty results -auto empty = QueryResults::Empty(num_queries); -``` - ---- - -## ItemID - -`ItemID` is a type alias for unique identifiers used throughout CyborgDB. - -```cpp -using ItemID = std::string; -``` - -`ItemID` is used to uniquely identify vectors and items within an encrypted index. Currently implemented as `std::string` for flexibility and human-readable identifiers. - ---- - -### `Item` - -`Item` struct holds the individual results from a `Get` operation, including the requested fields. - -```cpp -struct Item { - const std::string id; // Item ID - const std::vector vector; // Vector embedding - const std::vector contents; // Decrypted contents - const std::string metadata; // Metadata (JSON string) -}; -``` - ---- - -## ResultFields - -`ResultFields` enum specifies which fields to include in query results. - -```cpp -enum class ResultFields { - kDistance, // Include distance scores in query results - kMetadata // Include metadata in query results -}; -``` - ---- - -### ItemFields - -`ItemFields` enum defines the fields that can be requested for an `Item` object. - -```cpp -enum class ItemFields { - kVector, // Include vector in returned items - kMetadata, // Include metadata in returned items - kContents // Include content data in returned items -}; -``` - -By default, `ids` are always included in the returned items. - ---- - -## KeyContext - -`KeyContext` carries the key material for a data operation. It holds the 32-byte index KEK and, for RBAC deployments, a 16-byte user identifier. A bare 32-byte index key (the `index_key`) implicitly converts to a `KeyContext`, so most callers pass the key directly; RBAC users construct one explicitly with their own `user_kek` and `user_id`. - -```cpp -// Root access: a bare 32-byte index KEK converts implicitly. -std::array index_key = {/* ... */}; -index->Query(q, cyborg::QueryParams{}, index_key); - -// RBAC user access: pass the user's 32-byte KEK and 16-byte user ID. -std::array user_kek = {/* ... */}; -std::array user_id = {/* ... */}; -index->Query(q, cyborg::QueryParams{}, cyborg::KeyContext{user_kek, user_id}); -``` - -| Field | Type | Description | -|-----------|--------------------------|-----------------------| -| `kek` | `std::array` | The 32-byte key for the operation — the root `index_key`, or an RBAC user's `user_kek`. | -| `user_id` | `std::array` | 16-byte RBAC user identifier. Omit for root access. | - -Operations that require the root index KEK (such as `DeleteIndex` and user management) reject a per-user `KeyContext`. See [Managing Users](./encrypted-index/manage-users) for RBAC details. - ---- - -## KMSBlob - -`KMSBlob` describes how an index's Key-Encryption-Key (KEK) is wrapped by an external KMS. It is persisted per index via the module-level KMS functions (see the [KMS](./kms) reference). This is primarily for service-layer deployments; embedded SDK users supplying their own KEK can ignore it. - -```cpp -struct KMSBlob { - std::string kms_name; // Logical KMS name - std::string provider; // "aws" | "aws-kms" | "none" - std::string key_id; // KMS key identifier - std::string region; // KMS region - std::vector wrapped_kek; // Wrapped KEK bytes - uint32_t version = 0; // Envelope version - int64_t created_at = 0; // Unix epoch seconds -}; -``` diff --git a/versions/next/embedded/guides/advanced/access-control.mdx b/versions/next/embedded/guides/advanced/access-control.mdx index eaeb907..a764132 100644 --- a/versions/next/embedded/guides/advanced/access-control.mdx +++ b/versions/next/embedded/guides/advanced/access-control.mdx @@ -13,7 +13,7 @@ CyborgDB supports per-user access control (RBAC) on an encrypted index. Instead ## Minting User Keys -The root holder calls `create_user_keys` (Python) / `CreateUserKeys` (C++), supplying the new user's 16-byte identifier, their 32-byte per-user KEK, the permissions to grant, and the **root** index KEK as the admin gate. +The root holder calls `create_user_keys`, supplying the new user's 16-byte identifier, their 32-byte per-user KEK, the permissions to grant, and the **root** index KEK as the admin gate. ```python Python icon="python" @@ -36,31 +36,6 @@ writer_kek = secrets.token_bytes(32) index.create_user_keys(writer_id, writer_kek, permissions=["read", "write"], index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include -#include - -cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); - -// Root index KEK — the admin gate for user management -std::array index_key; -RAND_bytes(index_key.data(), index_key.size()); -auto index = client.CreateIndex("shared_index", index_key); - -// Mint a read-only user -std::array reader_id; RAND_bytes(reader_id.data(), reader_id.size()); -std::array reader_kek; RAND_bytes(reader_kek.data(), reader_kek.size()); -index->CreateUserKeys(reader_id, index_key, reader_kek, - /*grant_read=*/true, /*grant_write=*/false); - -// Mint a read/write user -std::array writer_id; RAND_bytes(writer_id.data(), writer_id.size()); -std::array writer_kek; RAND_bytes(writer_kek.data(), writer_kek.size()); -index->CreateUserKeys(writer_id, index_key, writer_kek, - /*grant_read=*/true, /*grant_write=*/true); -``` You must securely deliver each user's `user_id` and `user_kek` to that user out-of-band — together they are that user's credentials for the index. @@ -69,7 +44,7 @@ You must securely deliver each user's `user_id` and `user_kek` to that user out- ## Loading the Index as a User -A user loads the index with their own KEK as the `index_key` and passes their `user_id`. From then on, every operation is gated to the permissions they hold — a read-only user's `upsert`/`delete` is rejected, raising `RuntimeError` in Python (`std::runtime_error` in C++). +A user loads the index with their own KEK as the `index_key` and passes their `user_id`. From then on, every operation is gated to the permissions they hold — a read-only user's `upsert`/`delete` is rejected, raising `RuntimeError`. ```python Python icon="python" @@ -88,21 +63,9 @@ results = user_index.query( # user_index.upsert([...], index_key=reader_kek, user_id=reader_id) # raises RuntimeError — no write wrap ``` -```cpp C++ icon="brackets-curly" -// The read-only user loads the index with their own key + user_id -auto user_index = client.LoadIndex("shared_index", reader_kek, nullptr, reader_id); - -// Per-op, build a KeyContext from the user's KEK + user_id -cyborg::KeyContext user_ctx{reader_kek, reader_id}; - -cyborg::Array2D q{{0.1f, 0.2f, 0.3f, 0.4f}}; -cyborg::QueryResults results = user_index->Query(q, cyborg::QueryParams{5}, user_ctx); - -// Writes throw std::runtime_error for a read-only user (no write wrap) -``` -In Python, every data operation requires the `index_key=` keyword argument; an RBAC user passes their KEK as `index_key=` together with `user_id=`. In C++, every key-bearing method takes a trailing `KeyContext`; construct it as `cyborg::KeyContext{user_kek, user_id}` for an RBAC user, or pass a bare KEK for the root holder. +Every data operation requires the `index_key=` keyword argument; an RBAC user passes their KEK as `index_key=` together with `user_id=`. --- @@ -120,15 +83,6 @@ for u in index.list_user_keys(index_key=index_key): index.delete_user_keys(reader_id, index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -// List all users (root-gated) -for (const auto& u : index->ListUserKeys(index_key)) { - std::cout << "read=" << u.has_read << " write=" << u.has_write << "\n"; -} - -// Revoke a user (idempotent, root-gated) -index->DeleteUserKeys(reader_id, index_key); -``` --- @@ -141,7 +95,4 @@ For full signatures and exceptions, refer to the API Reference: API reference for `create_user_keys` / `list_user_keys` / `delete_user_keys` in Python - - API reference for `CreateUserKeys` / `ListUserKeys` / `DeleteUserKeys` in C++ - diff --git a/versions/next/embedded/guides/advanced/configure-index.mdx b/versions/next/embedded/guides/advanced/configure-index.mdx index fc6df5a..eac0555 100644 --- a/versions/next/embedded/guides/advanced/configure-index.mdx +++ b/versions/next/embedded/guides/advanced/configure-index.mdx @@ -31,20 +31,6 @@ index_key = secrets.token_bytes(32) # 32-byte index KEK index = client.create_index("test_index", index_key) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include -#include - -cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); - -std::array index_key; -RAND_bytes(index_key.data(), index_key.size()); // 32-byte index KEK - -// Create a DiskIVF index with defaults (dimension auto-detected on first upsert) -auto index = client.CreateIndex("test_index", index_key); -``` ### Creation Parameters @@ -52,7 +38,7 @@ auto index = client.CreateIndex("test_index", index_key); You can override the defaults at creation time: - `dimension`: vector dimensionality. Optional — auto-detected from the first upsert, or derived from `embedding_model` if provided. -- `storage_precision`: the on-disk dtype used for the rerank vectors. `float32` (default) gives the highest recall; `float16` roughly **halves disk footprint** with a slight precision loss. Acceptable values are `numpy.float32` / `numpy.float16` (or the strings `"float32"` / `"float16"`) in Python, and `StoragePrecision::Float32` / `StoragePrecision::Float16` in C++. +- `storage_precision`: the on-disk dtype used for the rerank vectors. `float32` (default) gives the highest recall; `float16` roughly **halves disk footprint** with a slight precision loss. Acceptable values are `numpy.float32` / `numpy.float16` (or the strings `"float32"` / `"float16"`) in Python. - `embedding_model`: an optional built-in model name that enables automatic text embedding (no `sentence-transformers`/`torch` install required) and overwrites the dimension. Call `cyborgdb.supported_embedding_models()` for accepted names. - `metric`: the distance metric — `"euclidean"` (default), `"cosine"`, or `"squared_euclidean"`. @@ -74,20 +60,6 @@ index = client.create_index( ) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" - -// Build a DiskIVF config: dimension 768, float16 rerank storage -cyborg::IndexDiskIVF index_config( - /*dimension=*/768, - /*embedding_model=*/"", - cyborg::StoragePrecision::Float16); // halve disk footprint vs. Float32 - -// Create the index with a cosine metric -auto index = client.CreateIndex( - "test_index", index_key, index_config, cyborg::DistanceMetric::Cosine); -``` Use `float16` storage precision when disk footprint matters more than the last fraction of a percent of recall. For most workloads the recall difference is negligible. @@ -116,17 +88,6 @@ index.train( ) ``` -```cpp C++ icon="brackets-curly" -// Train the index with custom clustering parameters -cyborg::TrainingConfig training_config( - /*n_lists=*/4096, - /*batch_size=*/0, // auto - /*max_iters=*/100, - /*tolerance=*/1e-6, - /*max_memory=*/0); // no limit - -index->TrainIndex(training_config, index_key); -``` For more on the training lifecycle, see [Training an Encrypted Index](../encrypted-indexes/train-index). @@ -152,19 +113,6 @@ results = index.query( ) ``` -```cpp C++ icon="brackets-curly" -// Tune recall vs. latency at query time -cyborg::Array2D query_vectors{{0.5, 0.9, 0.2, 0.7}}; -cyborg::QueryParams query_params( - /*top_k=*/10, - /*n_probes=*/32, - /*filters=*/"", - /*include=*/{}, - /*greedy=*/false, - /*rerank_mult=*/10); - -cyborg::QueryResults results = index->Query(query_vectors, query_params, index_key); -``` --- @@ -184,11 +132,6 @@ index = client.create_index( ) ``` -```cpp C++ icon="brackets-curly" -// Existing setup ... - -auto index = client.CreateIndex("index_name", index_key, cyborg::DistanceMetric::Cosine); -``` The currently supported distance metrics are: @@ -207,7 +150,4 @@ For more information on configuring an encrypted index, refer to the API Referen API reference for `StoragePrecision` and index types in Python - - API reference for `IndexDiskIVF` and `StoragePrecision` in C++ - diff --git a/versions/next/embedded/guides/advanced/kms.mdx b/versions/next/embedded/guides/advanced/kms.mdx index c6c4253..ea1fb74 100644 --- a/versions/next/embedded/guides/advanced/kms.mdx +++ b/versions/next/embedded/guides/advanced/kms.mdx @@ -50,26 +50,6 @@ cyborgdb.create_index_kms(config, "my_index", blob) cyborgdb.push_index_kms(config, "my_index", blob) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/index_kms.hpp" - -// Where the index keystore lives -cyborg::StorageConfig config = cyborg::StorageConfig::Disk("/tmp/cyborgdb"); - -cyborg::KMSBlob blob; -blob.kms_name = "prod-kms"; -blob.provider = "aws-kms"; -blob.key_id = "arn:aws:kms:us-east-1:123456789012:key/abcd-..."; -blob.region = "us-east-1"; -blob.wrapped_kek = {/* KEK wrapped by the external KMS */}; -blob.version = 1; - -// Strict insert (throws if an envelope already exists for this index) -cyborg::CreateIndexKMS(config, "my_index", blob); - -// Idempotent upsert (creates or overwrites) -cyborg::PushIndexKMS(config, "my_index", blob); -``` --- @@ -92,16 +72,6 @@ print(blob.provider, blob.key_id, blob.version) cyborgdb.delete_index_kms(config, "my_index") ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/index_kms.hpp" - -// Read the envelope back -cyborg::KMSBlob blob = cyborg::GetIndexKMS(config, "my_index"); -// ... unwrap blob.wrapped_kek with your KMS to recover the 32-byte KEK ... - -// Delete the envelope (idempotent) -cyborg::DeleteIndexKMS(config, "my_index"); -``` For embedded deployments where you already hold the KEK, you don't need the KMS envelope at all — pass the KEK directly to `create_index` / `load_index`. See [Managing Encryption Keys](./managing-keys) and [Access Control (RBAC)](./access-control) for per-user keys. @@ -116,7 +86,4 @@ For full signatures and the `KMSBlob` type, refer to the API Reference: API reference for the KMS module functions and `KMSBlob` in Python - - API reference for the KMS module functions and `KMSBlob` in C++ - diff --git a/versions/next/embedded/guides/advanced/managing-keys.mdx b/versions/next/embedded/guides/advanced/managing-keys.mdx index c9b6e0c..aec390a 100644 --- a/versions/next/embedded/guides/advanced/managing-keys.mdx +++ b/versions/next/embedded/guides/advanced/managing-keys.mdx @@ -35,13 +35,6 @@ The `index_key` must be **exactly 32 raw bytes**. The simplest way to generate o index_key = secrets.token_bytes(32) # 32-byte index KEK ``` - ```cpp C++ icon="brackets-curly" - #include - #include - - std::array index_key; - RAND_bytes(index_key.data(), index_key.size()); // 32-byte index KEK - ``` Alternatively, generate a key on the command line and persist the raw bytes to a file: @@ -71,16 +64,6 @@ The `index_key` must be **exactly 32 raw bytes**. The simplest way to generate o index = client.create_index("test_index", index_key) ``` - ```cpp C++ icon="brackets-curly" - #include "cyborgdb_core/client.hpp" - #include "cyborgdb_core/encrypted_index.hpp" - - // In-memory backing store (use StorageConfig::Disk(path) or ::S3(bucket) to persist) - cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); - - // index_key holds the 32 raw bytes generated above - auto index = client.CreateIndex("test_index", index_key); - ``` diff --git a/versions/next/embedded/guides/data-operations/add-items.mdx b/versions/next/embedded/guides/data-operations/add-items.mdx index edaab1c..631ba58 100644 --- a/versions/next/embedded/guides/data-operations/add-items.mdx +++ b/versions/next/embedded/guides/data-operations/add-items.mdx @@ -20,25 +20,15 @@ items = [ index.upsert(items, index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -// Items to add to the encrypted index -std::vector ids = {"item_1", "item_2", "item_3"}; -cyborg::Array2D vectors{{0.1, 0.2, 0.3, 0.4}, {0.5, 0.6, 0.7, 0.8}, {0.9, 0.10, 0.11, 0.12}}; - -// Add items to the encrypted index (index_key implicitly builds the KeyContext) -index->Upsert(ids, vectors, /*contents=*/{}, /*metadata=*/{}, index_key); -``` -For more info on `Array2D` in C++, see the [API Reference](../../cpp/types#array2d). - -In Python, every data operation takes `index_key` as a keyword argument — the client does not cache the key across calls. In C++ every key-bearing method takes a trailing `KeyContext`, which constructs implicitly from a bare 32-byte KEK — so `index->Upsert(..., index_key)` is the common path. You can also pass `user_id=` (Python) / a scoped `KeyContext` (C++) for an RBAC user. +Every data operation takes `index_key` as a keyword argument — the client does not cache the key across calls. You can also pass `user_id=` for an RBAC user. ## Adding Items with Contents It's also possible to store item contents alongside vectors. To do this, include `contents` to the `upsert()` call. -For Python, the contents field accepts both strings and bytes. For C++, the contents field only accepts bytes. All contents are encoded to bytes and encrypted before storage using the index key, and will be returned as bytes when retrieved with `get()`. +The contents field accepts both strings and bytes. All contents are encoded to bytes and encrypted before storage using the index key, and will be returned as bytes when retrieved with `get()`. ```python Python icon="python" @@ -53,19 +43,6 @@ items = [ index.upsert(items, index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -// Items to add to the encrypted index -std::vector ids = {"item_1", "item_2", "item_3"}; -cyborg::Array2D vectors{{0.1, 0.2, 0.3, 0.4}, {0.5, 0.6, 0.7, 0.8}, {0.9, 0.10, 0.11, 0.12}}; -std::vector> contents = { - std::vector{'H', 'e', 'l', 'l', 'o', '!'}, - std::vector{'W', 'o', 'r', 'l', 'd', '!'}, - std::vector{'C', 'y', 'b', 'o', 'r', 'g', '!'} -}; - -// Add items to the encrypted index -index->Upsert(ids, vectors, contents, /*metadata=*/{}, index_key); -``` ## Adding Items with Metadata @@ -87,19 +64,6 @@ items = [ index.upsert(items, index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -// Items to add to the encrypted index -std::vector ids = {"item_1", "item_2", "item_3"}; -cyborg::Array2D vectors{{0.1, 0.2, 0.3, 0.4}, {0.5, 0.6, 0.7, 0.8}, {0.9, 0.10, 0.11, 0.12}}; -std::vector metadata = { - R"({"name": "Alice", "age": 30})", - R"({"name": "Bob", "age": 40})", - R"({"name": "Charlie", "age": 50})" -}; - -// Add items to the encrypted index -index->Upsert(ids, vectors, /*contents=*/{}, metadata, index_key); -``` For more info on metadata storage and filtering, see [Metadata Filtering](./metadata-filtering). @@ -139,7 +103,4 @@ For more information on adding items to an encrypted index, refer to the API ref API reference for `upsert()` in Python - - API reference for `Upsert()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/data-operations/delete-items.mdx b/versions/next/embedded/guides/data-operations/delete-items.mdx index c0bfd4b..80733bc 100644 --- a/versions/next/embedded/guides/data-operations/delete-items.mdx +++ b/versions/next/embedded/guides/data-operations/delete-items.mdx @@ -12,9 +12,6 @@ You can delete items from an encrypted index using `delete()`: index.delete(["item1", "item2"], index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -index->Delete({"item1", "item2"}, index_key); -``` This operation is irreversible. Once you delete an item, you cannot recover it. @@ -27,7 +24,4 @@ For more information on deleting items from an encrypted index, refer to the API API reference for `delete()` in Python - - API reference for `Delete()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/data-operations/get-items.mdx b/versions/next/embedded/guides/data-operations/get-items.mdx index ff61590..c2cd188 100644 --- a/versions/next/embedded/guides/data-operations/get-items.mdx +++ b/versions/next/embedded/guides/data-operations/get-items.mdx @@ -7,7 +7,7 @@ noindex: true If you have added items to the index with the `contents` field, you can retrieve them via `get()`. -The `contents` field is always returned as bytes in Python or `std::vector` in C++, regardless of whether it was originally stored as a string or bytes. All contents are encoded to bytes and encrypted before storage. +The `contents` field is always returned as bytes, regardless of whether it was originally stored as a string or bytes. All contents are encoded to bytes and encrypted before storage. ```python Python icon="python" @@ -24,19 +24,6 @@ print(items) # {"id": "item_11", "contents": b"Hello, Cyborg!", "metadata": {"type": "md"}}] ``` -```cpp C++ icon="brackets-curly" -// IDs of items to retrieve -std::vector ids = {"item_20", "item_11"}; -std::vector include = {cyborg::ItemFields::kContents, cyborg::ItemFields::kMetadata}; - -// Retrieve items from the encrypted index (index_key implicitly builds the KeyContext) -std::vector items = index->Get(ids, include, index_key); - -// Print the item fields -for (const auto& item : items) { - std::cout << "ID: " << item.id << ", Contents: " << std::string(item.contents.begin(), item.contents.end()) << ", Metadata: " << item.metadata << std::endl; -} -``` ## API Reference @@ -47,7 +34,4 @@ For more information on getting items from an encrypted index, refer to the API API reference for `get()` in Python - - API reference for `Get()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/data-operations/metadata-filtering.mdx b/versions/next/embedded/guides/data-operations/metadata-filtering.mdx index fc05239..76fed15 100644 --- a/versions/next/embedded/guides/data-operations/metadata-filtering.mdx +++ b/versions/next/embedded/guides/data-operations/metadata-filtering.mdx @@ -32,15 +32,6 @@ data = [ index.upsert(data, index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -// Example data -std::vector ids = {"item_1", "item_2"}; -cyborg::Array2D vectors{{0.1, 0.1, 0.1, 0.1}, {0.2, 0.2, 0.2, 0.2}}; -std::vector metadata = {R"({"category": "dog"})", R"({"category": "cat"})"}; - -// Upsert data with metadata -index->Upsert(ids, vectors, /*contents=*/{}, metadata, index_key); -``` This metadata will be encrypted and stored in the index. @@ -68,21 +59,6 @@ filters = { results = index.query(query_vectors=query_vectors, top_k=top_k, filters=filters, index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -// Example query -cyborg::Array2D query_vectors{{0.5, 0.9, 0.2, 0.7}}; -int top_k = 10; - -// Example filters -std::string filters = R"({"category": {"$in": ["dog", "cat"]}})"; - -// Create QueryParams -cyborg::QueryParams query_params(top_k); -query_params.filters = filters; - -// Perform query -cyborg::QueryResults results = index->Query(query_vectors, query_params, index_key); -``` ## Metadata Indexing diff --git a/versions/next/embedded/guides/data-operations/query.mdx b/versions/next/embedded/guides/data-operations/query.mdx index cb995e7..182cce6 100644 --- a/versions/next/embedded/guides/data-operations/query.mdx +++ b/versions/next/embedded/guides/data-operations/query.mdx @@ -25,21 +25,6 @@ for r in results: print(r["id"]) ``` -```cpp C++ icon="brackets-curly" -// Example query -cyborg::Array2D query_vectors{{0.5, 0.9, 0.2, 0.7}}; -size_t top_k = 10; - -// Perform query (index_key implicitly builds the KeyContext) -cyborg::QueryParams query_params(top_k); -cyborg::QueryResults results = index->Query(query_vectors, query_params, index_key); - -// Print the results -auto view = results[0]; -for (uint32_t i = 0; i < view.num_results; ++i) { - std::cout << "ID: " << view.ids[i] << ", Distance: " << view.distances[i] << std::endl; -} -``` ## Query Parameters @@ -81,26 +66,6 @@ for r in results: print(f"ID: {r['id']}, Distance: {r['distance']}, Metadata: {r['metadata']}") ``` -```cpp C++ icon="brackets-curly" -// Example query -cyborg::Array2D query_vectors{{0.5, 0.9, 0.2, 0.7}}; -size_t top_k = 10; -size_t n_probes = 5; -bool greedy = false; -std::string filters = R"({"age": {"$gt": 18}})"; -std::vector include = {cyborg::ResultFields::kDistance, cyborg::ResultFields::kMetadata}; -size_t rerank_mult = 10; - -// Perform query (params + key are required) -cyborg::QueryParams query_params(top_k, n_probes, filters, include, greedy, rerank_mult); -cyborg::QueryResults results = index->Query(query_vectors, query_params, index_key); - -// Print the results -auto view = results[0]; -for (uint32_t i = 0; i < view.num_results; ++i) { - std::cout << "ID: " << view.ids[i] << ", Distance: " << view.distances[i] << ", Metadata: " << view.metadata[i] << std::endl; -} -``` ## Batched Queries @@ -117,15 +82,6 @@ top_k = 10 results = index.query(query_vectors=query_vectors, top_k=top_k, index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -// Example batch query -cyborg::Array2D query_vectors{{0.5, 0.9, 0.2, 0.7}, {0.1, 0.3, 0.8, 0.6}}; -size_t top_k = 10; - -// Perform batch query -cyborg::QueryParams query_params(top_k); -cyborg::QueryResults results = index->Query(query_vectors, query_params, index_key); -``` ## Querying with Metadata Filters @@ -185,7 +141,4 @@ For more information on querying encrypted indexes, refer to the API reference: API reference for `query()` in Python - - API reference for `Query()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/encrypted-indexes/create-client.mdx b/versions/next/embedded/guides/encrypted-indexes/create-client.mdx index f4fd082..ff7a45e 100644 --- a/versions/next/embedded/guides/encrypted-indexes/create-client.mdx +++ b/versions/next/embedded/guides/encrypted-indexes/create-client.mdx @@ -18,23 +18,12 @@ import cyborgdb_core as cyborgdb client = cyborgdb.Client(storage_config=cyborgdb.StorageConfig.disk("/tmp/cyborgdb-dev")) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include - -// In-memory backing store (ephemeral, for development/tests). -// Pass "" for api_key to run in free-tier mode. -cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); -``` To lift the free-tier cap, pass an API key as the first argument (`cyborgdb.Client(api_key, storage_config=...)`). Get a key from the [CyborgDB Admin Dashboard](https://cyborgdb.co); for more info, follow [this guide](../../../intro/get-api-key). Bear in mind that all contents stored in the backing store are end-to-end encrypted, meaning that **no index contents are stored in plaintext**. -The C++ `cyborg::Client` is non-copyable and non-movable (it owns the keystore handles). Construct it in place, or hold it via `std::unique_ptr` if you need to move ownership. - ## Persistent Backing Stores `StorageConfig` exposes two static factories: `disk(path)` (local disk) and `s3(bucket)` (AWS S3 or S3-compatible). @@ -54,20 +43,9 @@ client = cyborgdb.Client(storage_config=cyborgdb.StorageConfig.s3( )) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" - -// Local persistent storage on disk (pass "" for api_key to run in free-tier mode) -cyborg::Client disk_client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); - -// Or AWS S3 / S3-compatible (MinIO, etc.) -cyborg::S3Options s3_opts; -s3_opts.region = "us-east-1"; -cyborg::Client s3_client("", cyborg::StorageConfig::S3("my-bucket", s3_opts), 0, cyborg::kNone); -``` -Omit `credentials=` (Python) / leave `credentials` unset in `S3Options` (C++) to use the AWS default credential provider chain (environment variables, `~/.aws/credentials`, EC2 instance profile, EKS IRSA). The disk store also accepts cache options (`cache_vectors`, `cache_metadata`, `cache_ids`) to keep hot data in memory. +Omit `credentials=` to use the AWS default credential provider chain (environment variables, `~/.aws/credentials`, EC2 instance profile, EKS IRSA). The disk store also accepts cache options (`cache_vectors`, `cache_metadata`, `cache_ids`) to keep hot data in memory. ## Setting Device Configurations @@ -92,14 +70,6 @@ client = cyborgdb.Client( ) ``` -```cpp C++ icon="brackets-curly" -// ... existing setup - -// Enable GPU for specific operations (upsert and train) -cyborg::GPUConfig gpu_config = cyborg::kUpsert | cyborg::kTrain; - -cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 4, gpu_config); -``` `gpu_config` can only be set if running on a CUDA-enabled system with the CUDA driver installed. Use [`GPUConfig`](../../python/types#gpuconfig) to specify which operations (upsert, train, query) should use GPU acceleration. @@ -114,7 +84,4 @@ For more information on the `Client` class, refer to the API Reference: API reference for `Client` in Python - - API reference for `cyborg::Client` in C++ - diff --git a/versions/next/embedded/guides/encrypted-indexes/create-index.mdx b/versions/next/embedded/guides/encrypted-indexes/create-index.mdx index 27e9e58..6b132f4 100644 --- a/versions/next/embedded/guides/encrypted-indexes/create-index.mdx +++ b/versions/next/embedded/guides/encrypted-indexes/create-index.mdx @@ -22,24 +22,6 @@ index_key = secrets.token_bytes(32) index = client.create_index("my_index", index_key) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include -#include - -// Create a client (use the same backing store you'll load the index with) -cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); - -// Generate a 32-byte secure encryption key using OpenSSL -std::array index_key; -if (RAND_bytes(index_key.data(), index_key.size()) != 1) { - throw std::runtime_error("Failed to generate secure random key"); -} - -// Create an encrypted index (returns std::unique_ptr) -auto index = client.CreateIndex("my_index", index_key); -``` This creates a new encrypted index. The vector dimension is auto-detected from the first upsert (or derived from an `embedding_model`); you can also set it explicitly with the keyword-only `dimension` parameter. Other optional parameters include `metric` (`"euclidean"` default, `"cosine"`, or `"squared_euclidean"`) and `storage_precision`. `storage_precision` picks the on-disk format for the rerank vectors: `"float32"` (default) or `"float16"` to halve the footprint at a slight precision cost, or one of the TurboQuant tiers `"tq12"` / `"tq8"` / `"tq6"` / `"tq4"` (12/8/6/4 bits per dimension) to trade storage for a small recall/latency cost — `"tq4"` is the most aggressive (~8x smaller than `float32`, ~94% recall@100 on a 768-dimensional benchmark dataset). All tiers work with every metric; the choice is fixed at create time. @@ -50,7 +32,7 @@ This creates a new encrypted index. The vector dimension is auto-detected from t **Store `index_key` before you upsert any data.** CyborgDB has no key-recovery path — data encrypted with a lost key is unrecoverable. For anything past evaluation, keep the key in AWS Secrets Manager, Vault, or your KMS. See [Managing Keys](../advanced/managing-keys). -To lift the free-tier 1M-items-per-index cap, pass an API key as the first argument (`cyborgdb.Client(api_key, storage_config=...)` in Python; `cyborg::Client(api_key, ...)` in C++). Get a key from the [CyborgDB Admin Dashboard](https://cyborgdb.co). See [Get an API Key](../../../intro/get-api-key). +To lift the free-tier 1M-items-per-index cap, pass an API key as the first argument (`cyborgdb.Client(api_key, storage_config=...)`). Get a key from the [CyborgDB Admin Dashboard](https://cyborgdb.co). See [Get an API Key](../../../intro/get-api-key). ## Automatic Embedding Generation @@ -84,7 +66,4 @@ For more information on creating encrypted indexes, refer to the API reference: API reference for `create_index()` in Python - - API reference for `CreateIndex()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/encrypted-indexes/delete-index.mdx b/versions/next/embedded/guides/encrypted-indexes/delete-index.mdx index e9b5e29..bca1c62 100644 --- a/versions/next/embedded/guides/encrypted-indexes/delete-index.mdx +++ b/versions/next/embedded/guides/encrypted-indexes/delete-index.mdx @@ -27,24 +27,6 @@ index = client.load_index("my_index", index_key) index.delete_index(index_key=index_key) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include - -// Create a client (use the same backing store the index was created with) -cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); - -// Provide the index key used when creating the index -// Example key (32 bytes) -std::array index_key; - -// Load the encrypted index -auto index = client.LoadIndex("my_index", index_key); - -// Delete the index -index->DeleteIndex(index_key); -``` ## API Reference @@ -55,7 +37,4 @@ For more information on deleting an encrypted index, refer to the API reference: API reference for `delete_index()` in Python - - API reference for `DeleteIndex()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/encrypted-indexes/list-indexes.mdx b/versions/next/embedded/guides/encrypted-indexes/list-indexes.mdx index 389abe1..dbf0d46 100644 --- a/versions/next/embedded/guides/encrypted-indexes/list-indexes.mdx +++ b/versions/next/embedded/guides/encrypted-indexes/list-indexes.mdx @@ -20,18 +20,6 @@ print(indexes) # ["index_one", "index_two", "index_three"] ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" - -cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); - -auto indexes = client.ListIndexes(); - -// Print the indexes -for (const auto& index_name : indexes) { - std::cout << index_name << std::endl; -} -``` ## API Reference @@ -42,7 +30,4 @@ For more information on listing encrypted indexes, refer to the API reference: API reference for `list_indexes()` in Python - - API reference for `ListIndexes()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/encrypted-indexes/load-index.mdx b/versions/next/embedded/guides/encrypted-indexes/load-index.mdx index 9dcfd98..fb10aa4 100644 --- a/versions/next/embedded/guides/encrypted-indexes/load-index.mdx +++ b/versions/next/embedded/guides/encrypted-indexes/load-index.mdx @@ -22,24 +22,9 @@ index_key = bytes.fromhex("...") # the 64-char hex of your existing 32-byte ind index = client.load_index("my_index", index_key) ``` -```cpp C++ icon="brackets-curly" -#include "cyborgdb_core/client.hpp" -#include "cyborgdb_core/encrypted_index.hpp" -#include - -// Create a client (use the same backing store the index was created with) -cyborg::Client client("", cyborg::StorageConfig::Disk("/tmp/cyborgdb"), 0, cyborg::kNone); - -// Provide the index key used when creating the index -// Example key (32 bytes) -std::array index_key; - -// Load an encrypted index (returns std::unique_ptr) -auto index = client.LoadIndex("my_index", index_key); -``` -For role-based access control (RBAC), a per-user load passes the user's 16-byte identifier: `load_index("my_index", user_kek, user_id=user_id)` in Python, or `client.LoadIndex("my_index", user_kek, nullptr, user_id)` in C++. The supplied key then acts as that user's key-encryption key, and per-operation permissions are enforced. +For role-based access control (RBAC), a per-user load passes the user's 16-byte identifier: `load_index("my_index", user_kek, user_id=user_id)`. The supplied key then acts as that user's key-encryption key, and per-operation permissions are enforced. You will need to replace `index_key` with your own index encryption key. For production use, we recommend that you use an HSM or KMS solution. @@ -53,7 +38,4 @@ For more information on loading an encrypted index, refer to the API reference: API reference for `load_index()` in Python - - API reference for `LoadIndex()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/encrypted-indexes/train-index.mdx b/versions/next/embedded/guides/encrypted-indexes/train-index.mdx index 140980e..e6b555a 100644 --- a/versions/next/embedded/guides/encrypted-indexes/train-index.mdx +++ b/versions/next/embedded/guides/encrypted-indexes/train-index.mdx @@ -27,14 +27,6 @@ index.train( ) ``` -```cpp C++ icon="brackets-curly" -// Train the encrypted index with default configuration -index->TrainIndex(cyborg::TrainingConfig{}, index_key); - -// Train the index with specific configuration -cyborg::TrainingConfig config(1024, 0, 100, 1e-6, 0); -index->TrainIndex(config, index_key); -``` Training needs at least `10,000` vectors in the index (ingested via `upsert`) **and** at least `2 * n_lists` vectors. Below either threshold, `train` does not raise: it logs a warning and leaves the index untrained. Check `is_trained()` to confirm training took effect. @@ -74,7 +66,4 @@ For more information on training an encrypted index, refer to the API reference: API reference for `train()` in Python - - API reference for `TrainIndex()` in C++ - \ No newline at end of file diff --git a/versions/next/embedded/guides/intro/about.mdx b/versions/next/embedded/guides/intro/about.mdx index db29eda..36dd90d 100644 --- a/versions/next/embedded/guides/intro/about.mdx +++ b/versions/next/embedded/guides/intro/about.mdx @@ -5,7 +5,7 @@ mode: "wide" noindex: true --- -CyborgDB Embedded provides native libraries for Python and C++, enabling you to integrate confidential vector search directly into your applications. This deployment model gives you maximum control, performance, and security by running entirely within your infrastructure. +CyborgDB Embedded provides a native Python library, enabling you to integrate confidential vector search directly into your applications. This deployment model gives you maximum control, performance, and security by running entirely within your infrastructure. ## Why Choose CyborgDB Embedded? @@ -27,9 +27,9 @@ Deploy in air-gapped networks, edge devices, or custom infrastructure where exte - *Get running in 5 minutes with Python or C++* + *Get running in 5 minutes with Python* - Step-by-step guide covering both Python and C++ embedded library setup and usage. + Step-by-step guide covering embedded library setup and usage. @@ -48,7 +48,7 @@ CyborgDB Embedded integrates directly into your application process: ```mermaid flowchart TB subgraph Your_Applications [Your Application] - Library("CyborgDB Library (Python/C++)") + Library("CyborgDB Library (Python)") end Library --> Disk("Disk (local)") @@ -91,7 +91,7 @@ flowchart TB *Get hands-on with the quickstart guide* - Follow our comprehensive guide covering both Python and C++ setup + Follow our comprehensive guide covering Python setup *Explore detailed integration examples* diff --git a/versions/next/embedded/python/introduction.mdx b/versions/next/embedded/python/introduction.mdx index dc3d5da..2b2cd2e 100644 --- a/versions/next/embedded/python/introduction.mdx +++ b/versions/next/embedded/python/introduction.mdx @@ -11,7 +11,6 @@ The Python API for CyborgDB is split into two main classes within the `cyborgdb_ - `Client` – Handles configuration, index creation/loading, and listing available indexes. - `EncryptedIndex` – Provides data operations on a specific encrypted index such as upserting vectors, training the index, querying, and retrieving stored items. -This API is also available in [C++](../cpp/introduction). ## Module-level helpers diff --git a/versions/next/intro/deployment-models.mdx b/versions/next/intro/deployment-models.mdx index 277ab11..1953cac 100644 --- a/versions/next/intro/deployment-models.mdx +++ b/versions/next/intro/deployment-models.mdx @@ -67,7 +67,7 @@ Centralized deployment, monitoring, and maintenance. Update vector search capabi **Direct library integration** -Embed CyborgDB directly into your applications using Python or C++ libraries. This approach provides maximum control and performance by eliminating network overhead. +Embed CyborgDB directly into your applications using the Python library. This approach provides maximum control and performance by eliminating network overhead. Learn more about CyborgDB Embedded [here](../embedded/guides/intro/about). @@ -93,12 +93,7 @@ Access to low-level APIs for custom index configurations, memory management, and Embed confidential vector search directly in Python applications - - - *Native C++ integration* - - High-performance native integration for C++ applications - + --- diff --git a/versions/next/intro/using-docs.mdx b/versions/next/intro/using-docs.mdx index cefb009..0c72c2e 100644 --- a/versions/next/intro/using-docs.mdx +++ b/versions/next/intro/using-docs.mdx @@ -32,7 +32,7 @@ Our documentation is organized into four main sections: *Direct library integration* - - Python and C++ bindings + - Python library - Installation and setup guides - Performance tuning - Advanced configuration @@ -66,7 +66,7 @@ Our documentation is organized into four main sections: - - **Multiple Languages** - Python, JavaScript, TypeScript, Go, C++ + - **Multiple Languages** - Python, JavaScript, TypeScript, Go - **Copy Button** - One-click code copying - **Runnable Examples** - Complete working code @@ -163,10 +163,6 @@ import { Client } from 'cyborgdb'; const client = new Client({ baseUrl: 'http://localhost:8000', apiKey: 'your-key' }); ``` -```cpp C++ -#include -auto client = cyborgdb::Client("your-key"); -``` ### Runnable Examples diff --git a/versions/next/service/guides/intro/about.mdx b/versions/next/service/guides/intro/about.mdx index 4a2f7d4..112be8e 100644 --- a/versions/next/service/guides/intro/about.mdx +++ b/versions/next/service/guides/intro/about.mdx @@ -151,7 +151,7 @@ flowchart TB - **Maximum performance** - Need sub-millisecond latency with zero network overhead - - **Single-language environment** - Python or C++ applications exclusively + - **Single-language environment** - Python applications exclusively - **Air-gapped deployment** - No outbound network access (requires running without an API key) - **Custom integration** - Require deep integration with existing systems - **Resource optimization** - Want to eliminate network serialization overhead