Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,7 @@ endif()

set(CED_SRC
src/model_loader.cpp
src/gguf_check.cpp
src/ced_runner.cpp
src/fft.cpp
src/mel.cpp
Expand Down
22 changes: 22 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -163,6 +163,28 @@ ced_capi_free(ctx);

The per-PCM entry points take an arbitrary mono window, so a realtime consumer can call them on a sliding buffer for live recognition. There is also a struct-array variant (`ced_capi_classify_pcm`), a WAV-path variant (`ced_capi_classify_path_json`), and `ced_capi_classify_pcm_probs`, which writes every class score in class-index order (no sorting, no allocation) for callers that want the raw distribution. See `include/ced_capi.h` for the full API.

### Loading from memory

A model can also be loaded from bytes you already hold, with no file path and no temporary file, on Linux, macOS and Windows:

```c
// `buf` holds a complete GGUF file of `len` bytes (read from a network blob, an archive, a bundle, ...).
ced_ctx *ctx = ced_capi_load_from_memory(buf, len);
free(buf); // fine: the tensor data was copied during the call
if (!ctx) { fprintf(stderr, "%s\n", ced_capi_last_error(NULL)); return 1; }
```

The buffer is only read during the call, so you can free or overwrite it as soon as the call returns. For a moment the process holds the buffer and the model together. A truncated or corrupt buffer gives `NULL` and a message in `ced_capi_last_error(NULL)`; it never reads outside `[buf, buf + len)`. Like `ced_capi_load`, loads from several threads at once are safe.

If the model is one component of a larger GGUF (a bundle that holds several models), pass the whole file and the component prefix:

```c
// Every key and tensor of the component is stored as "<prefix><name>", e.g. "ced.encoder.init_bn.weight".
ced_ctx *ctx = ced_capi_load_from_memory_prefixed(bundle, bundle_len, "ced.");
```

Only the tensors under the prefix are copied, so no standalone copy of the component is needed. C++ users can call `ced::Ced::load_from_memory(data, size, prefix = "")`.

---

## LocalAI
Expand Down
30 changes: 30 additions & 0 deletions include/ced_capi.h
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
#ifndef CED_CAPI_H
#define CED_CAPI_H

#include <stddef.h>

#ifdef __cplusplus
extern "C" {
#endif
Expand All @@ -20,12 +22,40 @@ extern "C" {
typedef struct ced_ctx ced_ctx;

// ABI version. Bump on any breaking change below. v1: initial classify API.
// The *_load_from_memory* functions were added later and do not change the version.
int ced_capi_abi_version(void);

// Load a GGUF model. Returns an owning context or NULL on failure (on failure,
// ced_capi_last_error(NULL) holds the message). Release with ced_capi_free.
ced_ctx* ced_capi_load(const char* gguf_path);

// Load a GGUF model from memory instead of a file. `data` points to the bytes of
// a complete GGUF file of `size` bytes. Nothing is read from disk and no
// temporary file is made, on any platform.
//
// Ownership: the buffer is only read during this call. The loader copies the
// tensor data into its own memory, so the caller may free or overwrite `data`
// as soon as the call returns, whatever the result. While the call runs, the
// memory use peaks at about the buffer size plus the model size.
//
// A truncated or corrupt buffer is rejected: the call returns NULL and
// ced_capi_last_error(NULL) holds the reason. It never reads outside
// [data, data + size). Same threading rules as ced_capi_load: the error string
// is thread-local and loads from several threads at once are safe.
// Release the context with ced_capi_free.
ced_ctx* ced_capi_load_from_memory(const void* data, size_t size);

// As ced_capi_load_from_memory, for a model stored inside a larger GGUF such as
// a bundle. Every metadata key and every tensor name of the model is stored as
// `<prefix><name>` (for example prefix "ced." gives "ced.ced.depth" and
// "ced.encoder.init_bn.weight"). Only the tensors under the prefix are copied,
// so the caller can pass the bytes of the whole bundle and no standalone copy of
// the component is needed. `prefix` must not be NULL or empty (use
// ced_capi_load_from_memory for a standalone model). Same ownership rule: the
// buffer may be freed after the call. NULL, with the reason in
// ced_capi_last_error(NULL), if the prefix matches no tensors.
ced_ctx* ced_capi_load_from_memory_prefixed(const void* data, size_t size, const char* prefix);

// Free a context from ced_capi_load. Safe on NULL.
void ced_capi_free(ced_ctx* ctx);

Expand Down
20 changes: 18 additions & 2 deletions src/ced.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -24,9 +24,24 @@ struct Ced::Head {
};

bool Ced::load(const std::string& path) {
if (!loader_.load(path)) return false;
err_.clear();
if (!loader_.load(path)) { err_ = loader_.error(); return false; }
return finish_load();
}

bool Ced::load_from_memory(const void* data, size_t size, const std::string& prefix) {
err_.clear();
if (!loader_.load_from_memory(data, size, prefix)) { err_ = loader_.error(); return false; }
return finish_load();
}

// Everything after the I/O step; identical for every load path.
bool Ced::finish_load() {
backend_ = std::make_unique<Backend>();
if (!backend_->ok() || !loader_.realize_weights(*backend_)) return false;
if (!backend_->ok() || !loader_.realize_weights(*backend_)) {
err_ = "backend or weight setup failed";
return false;
}

const CedConfig& c = loader_.config();
const float* bw = loader_.host_f32("encoder.init_bn.weight");
Expand All @@ -35,6 +50,7 @@ bool Ced::load(const std::string& path) {
const float* bv = loader_.host_f32("encoder.init_bn.running_var");
if (!bw || !bb || !bm || !bv) {
std::fprintf(stderr, "ced: missing init_bn tensors\n");
err_ = "model is missing the encoder.init_bn tensors";
return false;
}
bn_scale_.resize(c.n_mels);
Expand Down
8 changes: 8 additions & 0 deletions src/ced.hpp
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,12 @@ namespace ced {
class Ced {
public:
bool load(const std::string& path);
// Load from a GGUF in memory (see ModelLoader::load_from_memory). `data` is
// not needed after the call returns. `prefix` selects a model inside a larger
// GGUF; leave it empty for a standalone model.
bool load_from_memory(const void* data, size_t size, const std::string& prefix = "");
// Why the last load failed ("" if it did not).
const std::string& load_error() const { return err_; }
const CedConfig& config() const { return loader_.config(); }
// Compute device the model runs on ("cpu", "CUDA0", "Vulkan0", ...).
const std::string& device_name() const { return backend_->device_name(); }
Expand Down Expand Up @@ -46,13 +52,15 @@ class Ced {
std::vector<float>& probs, int n_threads = 4);

private:
bool finish_load();
struct Embed;
struct Head;
Embed build_embed(ggml_context* ctx, std::vector<GraphInput>& inputs,
const std::vector<float>& input_values, int T) const;
Head build_blocks(ggml_context* ctx, ggml_tensor* tokens, int n_tokens) const;

ModelLoader loader_;
std::string err_;
std::unique_ptr<Backend> backend_;
// init_bn (BatchNorm2d, eval) folded into a per-mel scale/shift at load.
std::vector<float> bn_scale_, bn_shift_;
Expand Down
41 changes: 41 additions & 0 deletions src/ced_capi.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,8 @@

#include <algorithm>
#include <cstdlib>
#include <exception>
#include <new>
#include <cstring>
#include <numeric>
#include <string>
Expand Down Expand Up @@ -111,6 +113,45 @@ ced_ctx* ced_capi_load(const char* gguf_path) {
return reinterpret_cast<ced_ctx*>(c);
}

static ced_ctx* load_memory_impl(const void* data, size_t size, const std::string& prefix) {
if (!data || size == 0) {
g_load_error = "empty model buffer";
return nullptr;
}
auto* c = new (std::nothrow) CedContext();
if (!c) {
g_load_error = "out of memory";
return nullptr;
}
try {
if (!c->model.load_from_memory(data, size, prefix)) {
const std::string& why = c->model.load_error();
g_load_error = "failed to load model from memory: " +
why;
delete c;
return nullptr;
}
} catch (const std::exception& e) {
g_load_error = std::string("failed to load model from memory: ") + e.what();
delete c;
return nullptr;
}
g_load_error.clear();
return reinterpret_cast<ced_ctx*>(c);
}

ced_ctx* ced_capi_load_from_memory(const void* data, size_t size) {
return load_memory_impl(data, size, std::string());
}

ced_ctx* ced_capi_load_from_memory_prefixed(const void* data, size_t size, const char* prefix) {
if (!prefix || !*prefix) {
g_load_error = "null or empty prefix";
return nullptr;
}
return load_memory_impl(data, size, prefix);
}

void ced_capi_free(ced_ctx* ctx) { delete reinterpret_cast<CedContext*>(ctx); }

const char* ced_capi_last_error(const ced_ctx* ctx) {
Expand Down
100 changes: 100 additions & 0 deletions src/gguf_check.cpp
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
#include "gguf_check.hpp"

#include <cstdint>
#include <cstring>

namespace ced {

namespace {

struct Cursor {
const uint8_t* p;
size_t size;
size_t pos = 0;

bool have(uint64_t n) const { return n <= size - pos; }
bool skip(uint64_t n) {
if (!have(n)) return false;
pos += (size_t)n;
return true;
}
template <typename T>
bool read(T* out) {
if (!have(sizeof(T))) return false;
std::memcpy(out, p + pos, sizeof(T));
pos += sizeof(T);
return true;
}
};

// Size in bytes of a fixed-size GGUF value type, 0 for string/array/unknown.
size_t scalar_size(uint32_t t) {
switch (t) {
case 0: case 1: case 7: return 1; // u8, i8, bool
case 2: case 3: return 2; // u16, i16
case 4: case 5: case 6: return 4; // u32, i32, f32
case 10: case 11: case 12: return 8; // u64, i64, f64
default: return 0;
}
}

constexpr uint32_t kString = 8, kArray = 9;
constexpr uint64_t kMaxString = 1ull << 30; // same cap as ggml's reader

bool skip_string(Cursor& c) {
uint64_t n;
return c.read(&n) && n <= kMaxString && c.skip(n);
}

} // namespace

bool gguf_precheck(const void* data, size_t size, std::string* err) {
auto fail = [&](const char* m) {
if (err) *err = m;
return false;
};
Cursor c{static_cast<const uint8_t*>(data), size};
uint32_t version;
int64_t n_tensors, n_kv;
if (!c.have(4) || std::memcmp(c.p, "GGUF", 4) != 0) return fail("not a GGUF file (bad magic)");
c.pos = 4;
if (!c.read(&version) || (version != 2 && version != 3))
return fail("unsupported GGUF version");
if (!c.read(&n_tensors) || !c.read(&n_kv) || n_tensors < 0 || n_kv < 0)
return fail("GGUF header is truncated or has a negative count");
// Every key/value pair takes at least 13 bytes, so a count that cannot fit is
// corrupt (and must not drive a long loop).
if ((uint64_t)n_kv > (size - c.pos) / 13) return fail("GGUF metadata count exceeds the buffer size");

for (int64_t i = 0; i < n_kv; ++i) {
uint64_t klen;
if (!c.read(&klen) || klen > kMaxString || !c.have(klen))
return fail("GGUF metadata key is truncated");
if (klen == 0) return fail("GGUF metadata key has an empty name");
c.skip(klen);
uint32_t type;
if (!c.read(&type)) return fail("GGUF metadata is truncated");
if (type == kString) {
if (!skip_string(c)) return fail("GGUF string value is truncated");
} else if (type == kArray) {
uint32_t et;
uint64_t n;
if (!c.read(&et) || !c.read(&n)) return fail("GGUF array header is truncated");
if (et == kString) {
for (uint64_t j = 0; j < n; ++j)
if (!skip_string(c)) return fail("GGUF string array is truncated");
} else {
const size_t es = scalar_size(et);
if (es == 0) return fail("GGUF array has an unsupported element type");
if (n > size / es || !c.skip(n * es)) return fail("GGUF array is truncated");
}
} else {
const size_t s = scalar_size(type);
if (s == 0) return fail("GGUF metadata has an unknown value type");
if (!c.skip(s)) return fail("GGUF metadata is truncated");
}
}
return true;
}

} // namespace ced
20 changes: 20 additions & 0 deletions src/gguf_check.hpp
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
#pragma once
#include <cstddef>
#include <string>

namespace ced {

// Cheap structural check of a GGUF held in memory, run BEFORE the buffer is
// handed to ggml's reader. ggml validates sizes and bounds, but it still aborts
// the whole process (GGML_ASSERT) on a few well-formed-looking inputs, such as a
// metadata key with an empty name. A model buffer comes from the caller, so such
// input must become an error instead.
//
// Walks the header and the metadata key/value section with bounds checks on every
// read. Returns false and sets `err` on: a bad magic or version, an empty or
// oversized key, an unknown value type, a nested array, or a value that runs
// past the end of the buffer. It does not look at the tensor table: ggml reports
// those errors without aborting, and the loaders re-check every tensor range.
bool gguf_precheck(const void* data, size_t size, std::string* err);

} // namespace ced
Loading
Loading