Neon published Generate embeddings with Neon AI Gateway on 2 October 2026, bylined to Carlota Soto on the product team. The news is a second route on a gateway Neon already uses for chat. POST /v1/embeddings accepts the same OpenAI-compatible credential as chat completions. You point an OpenAI SDK at the branch’s gateway base URL, write the vectors into Lakebase Postgres, and query them with Lakebase Search. Neon says the gateway is billed at each lab’s published per-token price with no markup, on Databricks Foundation Model APIs.
The two models and the request shape
Neon says the endpoint takes one string or a batch of up to 150 strings and returns the usual OpenAI embedding object. Two models are listed at launch, with prices “as of Oct 2026”:
| Model | Dimensions | Normalized | Price |
|---|---|---|---|
| qwen3-embedding-0-6b | 1024, and Neon says this one is configurable | Yes | $0.02 per 1 million input tokens |
| gte-large-en | 1024 | No; Neon says use cosine distance | $0.13 per 1 million input tokens |
The sample calls qwen3-embedding-0-6b on the sentence “The quick brown fox jumps over the lazy dog.” and prints a length of 1024. It sets encoding_format to float, which the sample comment describes as required on OpenAI SDK v6 and harmless on v7 and later. The client uses NEON_AI_GATEWAY_TOKEN and NEON_AI_GATEWAY_BASE_URL. Neon says neon env pull writes both for the branch you are on, and that a Neon Function receives them without that step. A credential needs the ai_gateway:invoke scope.
Availability is narrower than the whole Neon free plan. The post says AI Gateway is on the Launch and Scale plans and is paid with prepaid credits. A free try, Neon says, is a matter of asking in Discord for credits. It suggests starting on qwen3-embedding-0-6b because it is the cheaper of the two, and using gte-large-en when you already have vectors from that model and need new ones to match.
Where the vectors go
The post treats Lakebase Search as the place the vectors land, not as the thing that shipped on 2 October. It describes lakebase_vector (an ANN index with pgvector types and operators) and lakebase_text (BM25 on tsvector). A pipeline sketch in the post is: a file lands in Object Storage, a Function trigger chunks it, the gateway embeds batches of up to 150, rows are upserted so a retry is idempotent, and a question is embedded with the same model, mixed with Reciprocal Rank Fusion, and sent to a chat model on the same gateway. Neon also says a branch copies the pipeline without copying storage, which is its pitch for re-embedding on a branch before you keep a new model.
Practical takeaway
If you already call Neon’s gateway for chat, embeddings are the same host, key, and /v1 prefix, on Launch or Scale, with two named models and a 150-input batch cap. Match the model to vectors you already stored. The $0.02 and $0.13 figures are Neon’s October 2026 table, so recheck them before you budget a backfill. The gateway does not provision itself on the free plan unless Neon has granted credits.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.