How to Install Needle 2 Locally on Windows (14MB AI)

You can install Needle 2 locally on Windows with one Python package. It is a specialist model for choosing tools and returning structured data, not a tiny replacement for ChatGPT. The Python package downloads a platform-specific engine on first use, then inference can run without a network connection.

Cactus Compute reports a 45-million-parameter model compressed into a 14 MB engine, with a full session using roughly 28 MB of RAM. These are project-reported figures; actual memory use and speed depend on the device and runtime.

The 14 MB figure describes the project-reported Needle engine, not the complete Python environment. Python, JAX, Flax and other dependencies use additional disk space.

Needle 2 local tool-calling AI routing structured tasks on Windows
Needle 2 is built for compact local tool routing and structured extraction, not general-purpose chat.

What Needle 2 is and is not

Needle 2 reads a request plus a list of functions or data fields. It can then select a function, fill its arguments or extract a record that matches a schema. Its output is constrained so the tool call itself is valid structured data.

Useful examples include:

  • routing a device command to the correct local function;
  • turning a short invoice or receipt into typed fields;
  • choosing from a large catalogue of narrow tools;
  • refusing an unrelated request when no declared tool can serve it.

It is not designed for open-ended knowledge questions or long-form writing. The official documentation says an unsupported request returns an empty call rather than a general chatbot answer. If you want a full local language model, the Qwen3.8-27B local guide covers a much larger general-purpose model. For local speech-to-text, use the NeMo-Speech.cpp transcription guide.

Needle 2 requirements

  • 64-bit Windows on x64 or ARM64 (the project does not document a minimum Windows release);
  • 64-bit Python 3.9 or later;
  • PowerShell or Windows Terminal;
  • an internet connection for the package and first engine download;
  • no dedicated GPU for the basic inference example.

The official model card lists Windows x64 and Windows ARM builds. The package record requires Python 3.9 or newer. A GPU is relevant to the optional fine-tuning extras, but it is not a prerequisite for this first local tool-calling example.

Step 1: Create a clean Windows environment

Open PowerShell in a folder you use for local AI projects, then create a separate virtual environment:

mkdir needle2-demo
cd needle2-demo
py -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade pip

Calling the environment’s Python executable directly avoids PowerShell activation-policy errors. If the py command is unavailable, install a current 64-bit Python release and enable the Python launcher during setup.

Step 2: Install Needle 2

.\.venv\Scripts\python.exe -m pip install --upgrade cactus-needle pydantic
.\.venv\Scripts\needle.exe --help

The current package checked for this guide is version 2.0.4, published on 14 August 2026. The project launched version 2.0.0 on 10 August and has already shipped several small package updates, so recording the installed version will make later troubleshooting easier:

.\.venv\Scripts\python.exe -m pip show cactus-needle

Step 3: Fetch the Windows engine

The model is baked into a native engine. The Python package fetches the correct engine once and caches it under the user’s cache directory.

.\.venv\Scripts\needle.exe fetch

This explicit fetch is useful because it separates network or platform problems from your Python code. The official API guide lists win_amd64 and win_arm64 engine tags. Do not download a similarly named binary from an unofficial repository.

Step 4: Connect a harmless local tool

Create a file named demo.py:

import needle

@needle.tool
def get_service_status(service: str):
    """Read the demo status of a named local service.

    Args:
        service: service name to look up
    """
    demo_status = {
        "web": "running",
        "backup": "idle",
        "database": "running",
    }
    return {
        "service": service,
        "status": demo_status.get(service.lower(), "unknown"),
    }

agent = needle.Needle(tools=[get_service_status])
result = agent.run("Check the web service")

print(result)

Run it with the virtual environment:

.\.venv\Scripts\python.exe demo.py

This example returns static demo data and cannot change the computer. That is deliberate. Needle’s run() loop can execute the Python functions you expose to it, so begin with read-only functions and inspect the returned call before connecting anything destructive, financial or security-sensitive.

Grammar-constrained JSON is not an authorization boundary. Validate tool names and every argument, enforce allowlists and permissions in ordinary code, rate-limit sensitive operations and require user confirmation for consequential actions.

Step 5: Inspect calls and confidence before acting

For higher-risk automation, use complete() and decide in your own code whether to execute the proposed function. The official response contract includes fields such as type, function_calls, confidence, error and error_code.

response = agent.complete("Check the database service")

print("Type:", response.get("type"))
print("Confidence:", response.get("confidence"))
print("Calls:", response.get("function_calls"))

Cactus Compute recommends setting a threshold appropriate to the product: act above it, and ask again or route to a larger model below it. The documentation does not define one universal safe number. A suitable threshold depends on the tool, the cost of a mistake and your own validation data.

Step 6: Extract a typed record

Needle treats structured extraction as a one-tool call. Add this to a separate file named extract_invoice.py:

import needle
from pydantic import BaseModel

class Invoice(BaseModel):
    vendor: str
    total: float
    due_date: str

text = "Invoice from Acme Corp, $1,200.00, due 2026-09-01"
invoice = needle.extract(text, Invoice)

if invoice is None:
    print("No supported invoice record was returned")
else:
    print(invoice.vendor, invoice.total, invoice.due_date)
.\.venv\Scripts\python.exe extract_invoice.py

Schema conformance does not prove that every extracted value is factually correct. Validate totals, dates and identifiers before putting them into accounting, ticketing or production systems.

Optional: open the local playground

.\.venv\Scripts\needle.exe playground

The official package documentation says this opens a local interface at http://127.0.0.1:7860. Keep it bound to localhost unless you have deliberately added authentication and network controls. A local development interface should not be exposed directly to the internet.

How offline setup works

Inference does not need the network after the engine is cached. The official API guide documents three useful controls:

  • needle fetch downloads the engine for the current machine and prints its path;
  • NEEDLE_LIB_PATH points Needle at a specific native library;
  • HF_HUB_OFFLINE=1 prevents a disconnected system from attempting a Hugging Face download.

Set offline mode only after copying or fetching the correct engine. Otherwise a missing file will fail immediately, which is useful for an air-gapped deployment, but confusing during the first setup.

Only model inference is offline. A Python tool you expose can still call a network service or write data elsewhere.

Troubleshooting Needle 2 on Windows

py is not recognised

Install 64-bit Python 3.9 or later from the official Python site and enable the launcher. Open a new terminal, then check:

py --version

PowerShell blocks Activate.ps1

You do not need to activate the environment. Use the full .\.venv\Scripts\python.exe and .\.venv\Scripts\needle.exe paths shown above.

The first run cannot download the engine

Run needle fetch separately so the error is easier to read. Check the proxy, firewall, system clock and access to Hugging Face. Do not enable HF_HUB_OFFLINE=1 until the engine is present.

A native library will not load

Confirm that Python and Windows use the same architecture, update cactus-needle, and fetch the engine again. The official project publishes separate Windows x64 and ARM builds. Avoid copying a Linux or macOS library into the Windows cache.

Needle returns an empty call

The official behaviour is to return [] when none of the declared tools can serve the request. Improve the function name, docstring and argument descriptions, or ask a request that genuinely matches the tool. Do not force every request into the nearest available action.

The wrong tool is selected

Make overlapping tools more distinct. Explain what each tool does, describe every argument and use Literal or needle.Field constraints where appropriate. With more than five declared tools, Needle uses its retrieval head to expose only the five highest-scoring tools for that turn; the official API supports tool_index_path to persist those embeddings.

Confidence disappears after fine-tuning

The API documentation says the calibrated confidence head applies to the base model. A locally fine-tuned model reports None because fine-tuning does not update that calibration. Validate a tuned model separately rather than assuming the base model’s threshold still applies.

Licence note

The official Hugging Face model card labels Needle 2 Apache-2.0. The source repository’s LICENSE file is MIT, while the current PyPI project page also displays Apache-2.0 metadata. That can reflect different licensing for model artefacts and source software, but the project does not explain the distinction on the pages checked for this guide. Review the exact model, binary and source files you redistribute, and ask the maintainer if the distinction matters to your use case.

Frequently asked questions

Is Needle 2 a normal chatbot?

No. It is designed for tool selection, device actions and structured extraction. An unrelated request normally produces an empty call instead of a general-knowledge answer.

Does Needle 2 run without a GPU?

The project publishes native Windows x64 and ARM64 builds, and its base inference instructions do not list a GPU requirement. GPU extras are for optional fine-tuning.

Is the model really 14 MB?

Cactus Compute reports that the CQ2 engine containing the model is 14 MB and a session uses roughly 28 MB of RAM. Actual runtime use varies by device, and the complete Python environment requires additional disk space.

Can Needle 2 execute functions?

The Python run() loop can execute decorated functions that you explicitly provide. Use harmless, read-only functions first. For sensitive actions, inspect the proposed call with complete(), apply your own permission checks and require confirmation where appropriate.

Can it work offline?

Yes, after the native engine has been downloaded or copied into place. The project documents a cache, an explicit library path and an offline environment variable. Only inference is offline; exposed Python tools may still use the network or write data.

Can I connect Needle 2 to MCP?

Needle consumes Python functions or JSON-schema tool descriptions. An MCP integration would need an adapter that turns MCP tool schemas into Needle’s tool format and executes approved calls. The official pages checked for this guide do not document a first-party MCP server, so do not assume drop-in compatibility. For a documented MCP-oriented local workflow, see the llama.cpp MCP setup guide.

Sources

Needle 2 is interesting because it narrows the job until a very small model becomes useful. The right first project is not autonomous computer control. It is one well-described, read-only tool or one schema whose output you can verify.

Leave a Reply

Scroll to Top

Discover more from Lachie's Lifestyle

Subscribe now to keep reading and get access to the full archive.

Continue reading