Skip to content

vgi-rpc C++

C++ implementation of the vgi_rpc framework — Apache Arrow IPC-based RPC for high-performance data services.

See Tailscale and trusted HTTP proxy identity for the peer-evidence adapters and their deployment trust boundaries.

Built by Query.Farm

Define RPC methods with typed C++20 handlers using Arrow schemas. The framework provides server dispatch with automatic parameter extraction and result serialization over four transports: stdin/stdout pipes, Unix domain sockets, TCP, and HTTP.

Key Features

  • Unary RPCs with typed parameter extraction via get<T>(name)
  • Producer streams for server-initiated batch data flows
  • Exchange streams for bidirectional batch processing
  • Client-directed logging at configurable levels
  • Introspection via the co-hosted vgi_rpc.Reflection.v1 protocol, which every client asks over the connection it already holds (list_protocols(), describe_protocol())
  • Error handling — exceptions automatically converted to protocol error responses
  • Builder pattern — fluent ServerBuilder API for registering methods
  • Access log — JSONL records per call, with the spec's size cap. Records describe a request by its shape (request_fields: parameter names and Arrow types; request_rows) and stream state tokens by size (request_state_bytes / response_state_bytes); no request, response or state value is ever logged, at any level

Transports

Transport Entry point Discovery line
Pipe (stdin/stdout) Server::run() —
Unix domain socket Server::serve_unix(path) UNIX:<path>
TCP (trusted networks; no auth or TLS) Server::serve_tcp(host, port) TCP:<host>:<port>
TCP behind a trusted L4 proxy Server::serve_tcp(host, port, TcpServerOptions) Required PROXY v2 + connection identity snapshot
HTTP Server::serve_http(HttpConfig) PORT:<port>

Pipe, Unix, and TCP share the same raw Arrow IPC framing and differ only in the socket they read and write. HTTP maps the same protocol onto stateless request/response pairs and carries the optional features below.

Shared memory

A side channel, not a transport of its own: it rides alongside a pipe or socket, which still carries control messages and small batches while a large batch is written into a POSIX segment and replaced on the wire by a zero-row pointer batch. Enable it with ServerBuilder::enable_transport_options(), which answers the __transport_options__ handshake — a worker that stays silent is read as "no shared memory" and peers stay inline.

Batches below VGI_RPC_SHM_MIN_BATCH_BYTES (128 KiB by default) stay inline, because the fixed cost of an allocation, a pointer round trip, and the peer's resolve and free only pays off once the copy it avoids is large enough.

HTTP features

Application and hosting byte ceilings remain optional. Every HTTP worker now advertises response-budget support so clients can require a bounded decoded response without guessing whether an older worker ignored the request header.

Feature Configuration
Capability discovery (GET/OPTIONS {prefix}/health) always on
Response caps, strict-fail max_response_bytes, hosting_max_response_bytes, max_externalized_response_bytes
Response batching target preferred_response_bytes
Request caps max_request_bytes, hosting_max_request_bytes
External locations (pointer batches) external_storage_url, externalize_threshold
Bounded request decoding and response negotiation (zstd + gzip) compression
CORS, including Cross-Origin-Resource-Policy cors_origin
Sticky sessions (AEAD-sealed tokens, TTL, drain) sticky, sticky_default_ttl, sticky_echo_headers
Standardized 401s with VGI-Auth-Reason reject_all
Proxy proof (HMAC-SHA256 proof-of-hop) proof_mode, proof_origin_id, proof_secrets

Stream state travels as two tokens split by lifetime — a call token minted once by /init and a cursor re-minted every turn — so a continuation does not re-serialize the fixed half of the call.

Clients send VGI-Accept-Max-Response-Bytes on every request. The effective decoded Arrow IPC limit is the minimum of the application, hosting, and client ceilings. A continuation can lower its init-time limit but cannot raise it. Oversize successful responses return an Arrow ResponseTooLargeError envelope with HTTP 200 and no continuation cursor. CallContext and OutputCollector expose the effective response_limit_bytes and the server-only preferred_response_bytes target to handlers.

External storage

A batch over externalize_threshold is uploaded and replaced on the wire by a zero-row pointer batch carrying a URL the client re-fetches. external_storage_url picks the backend by scheme:

Scheme Backend Build
http(s):// A service speaking the four-endpoint alloc/PUT/HEAD/GET contract always
s3://bucket/prefix AWS S3 via aws-sdk-cpp -DVGI_RPC_WITH_S3=ON
gs://bucket/prefix Google Cloud Storage via google-cloud-cpp -DVGI_RPC_WITH_GCS=ON

The cloud backends are off by default — both SDKs are long builds, and a deployment that externalises through its own HTTPS service needs neither. Turn them on together with the matching vcpkg manifest features:

cmake --preset default \
  -DVCPKG_MANIFEST_FEATURES="s3;gcs" \
  -DVGI_RPC_WITH_S3=ON -DVGI_RPC_WITH_GCS=ON

A URI naming a backend the binary was not built with is refused at startup, not on the first payload large enough to externalise.

What a pointer batch carries is always a pre-signed HTTPS URL, never a bucket path, so the client fetches it holding no cloud credentials and linking no SDK. signed_url_ttl_seconds bounds how long a leaked pointer stays usable. Uploaded objects are never deleted — set a lifecycle rule on the bucket.

Pre-published references

A result that is large and rarely changes — a whole catalog, say — need not be serialized and uploaded on every call. Publish it once with publish_external, cache the ExternalRef it returns, and answer later calls with Result::from_external_ref:

#include <vgi_rpc/external.h>

auto storage = vgi_rpc::make_external_storage({.uri = "https://objects.example/vgi"});
std::optional<vgi_rpc::ExternalRef> catalog_ref;  // guard with a mutex under HTTP

builder.add_unary("catalog", vgi_rpc::empty_schema(), catalog_schema,
    [&](const vgi_rpc::Request&, vgi_rpc::CallContext&) {
        if (!catalog_ref) {
            // A 1-row batch on the method's result schema.
            catalog_ref = vgi_rpc::publish_external(build_catalog_batch(), *storage, "zstd");
        }
        return vgi_rpc::Result::from_external_ref(*catalog_ref);
    });

The server writes the pointer batch (vgi_rpc.location, plus vgi_rpc.location.sha256 only when the ref has a digest) directly on every transport — pipe, Unix, TCP and HTTP — whether or not the server has storage configured and regardless of externalize_threshold. Nothing is serialized or uploaded during the call, the pointer never takes the shared-memory route, and it is not charged against max_externalized_response_bytes. Clients resolve it like any other pointer. publish_external serializes, hashes, compresses and uploads exactly as the per-call externalizer does; pass include_sha256 = false (or build ExternalRef(url) by hand) to omit the digest so clients skip the content check. Unary methods only.

You own the ref's cache and the object's lifecycle: a long-lived ref must not point at an object under the short-TTL lifecycle rule used for per-call uploads, and a pre-signed URL expires — re-sign or rebuild the ref before then. Only hand a ref to callers who are all entitled to the same content.

Three Method Types

Unary

A single request produces a single response. The client sends parameters, the server returns a result.

Client  ──  add(a=2, b=3)  ──▸  Server
Client  ◂──     5.0         ──  Server

Producer

The server pushes batches to the client until calling out.finish():

Client  ──  produce_n(count=3)  ──▸  Server
Client  ◂──  {index: [0]}       ──  Server
Client  ◂──  {index: [1]}       ──  Server
Client  ◂──  {index: [2]}       ──  Server
Client  ◂──    [finish]         ──  Server

Exchange

Lockstep bidirectional streaming — one request, one response, repeat:

Client  ──  exchange_scale(factor=2.5)  ──▸  Server
Client  ──    {value: [10.0]}           ──▸  Server
Client  ◂──   {value: [25.0]}          ──  Server
Client  ──    {value: [4.0]}            ──▸  Server
Client  ◂──   {value: [10.0]}          ──  Server
Client  ──      [close]                 ──▸  Server

Quick Example

examples/quick_example.cpp
// © Copyright 2025-2026, Query.Farm LLC - https://query.farm
// SPDX-License-Identifier: Apache-2.0

#include <vgi_rpc/server.h>
#include <vgi_rpc/request.h>
#include <vgi_rpc/result.h>
#include <vgi_rpc/arrow_utils.h>

#include <arrow/array/builder_primitive.h>
#include <arrow/type.h>

int main() {
    auto server =
        vgi_rpc::ServerBuilder()
            .add_unary(
                "add",
                arrow::schema({
                    arrow::field("a", arrow::float64()),
                    arrow::field("b", arrow::float64()),
                }),
                arrow::schema({arrow::field("result", arrow::float64())}),
                [](const vgi_rpc::Request& req, vgi_rpc::CallContext& /*ctx*/) {
                    double a = req.get<double>("a");
                    double b = req.get<double>("b");

                    arrow::DoubleBuilder builder;
                    VGI_RPC_THROW_NOT_OK(builder.Append(a + b));
                    auto array = vgi_rpc::unwrap(builder.Finish());

                    return vgi_rpc::Result::value(
                        arrow::schema({arrow::field("result", arrow::float64())}), {array});
                },
                "Add two numbers together.")
            .protocol("MyServer")
            .build();

    server->run();
}

Next Steps


vgi-rpc · Query.Farm

Embedded browser assets

HttpConfig::static_assets maps exact URL paths to HttpStaticAsset values containing a body and content type. Assets are mounted both at the HTTP prefix and at the root, use the same authentication and proxy-proof checks as RPC, and support HEAD and ETag revalidation. An optional json_body supplies the representation selected by ?format=json or an Accept: application/json request. Paths are literal and never read files from disk.