Skip to main content

Query & download activity logs

Activity originating at Nucliadb (like searches or questions) is stored on the activity log. You can either query it for instant paginated results, or request an asynchronous download of the full result set.

Downloads are asynchronous: you request a query, a file is prepared, and you either wait, poll for status, or get notified via email when it's ready.

Query Parameters

ParameterDescription
year_monthYear and month of logs to retrieve (e.g., 2024-02)
showFields to display in the output (id is always included)
filtersFilter criteria (see operators below)
paginationControl result size and cursor position

Filter Operators

OperatorDescription
eqEqual to
gt / geGreater than / Greater than or equal to
lt / leLess than / Less than or equal to
neNot equal to
isnullCheck for null (True/False)
likeSQL-like pattern (string fields only)
ilikeCase-insensitive SQL-like pattern (string fields only)
isinValue is in a given list
isnotinValue is not in a given list

Pagination

ParameterDescription
limitNumber of items to fetch
starting_afterFetch logs after a specific ID (ascending)
ending_beforeFetch logs before a specific ID (descending)

Available Fields

Common Fields (All Event Types)

id, date, user_id, user_type, client_type, total_duration, audit_metadata, resource_id, nuclia_tokens, token_details

SEARCH Events

Common fields + question, resources_count, filter, retrieval_rephrased_question, vectorset, security, min_score_bm25, min_score_semantic, result_per_page, retrieval_time

CHAT Events

Common fields + question, answer, rephrased_question, learning_id, retrieved_context, chat_history, feedback_good, feedback_comment, feedback_good_all, feedback_good_any, feedback, model, rag_strategies_names, rag_strategies, status, generative_answer_first_chunk_time, generative_reasoning_first_chunk_time, generative_answer_time, remi_scores, user_request, reasoning

ASK Events

All SEARCH fields + all CHAT fields.

Query Examples

CLI

nuclia kb logs query --type=ASK --query='{
"year_month": "2024-10",
"show": ["id", "date", "question", "answer", "feedback_good"],
"filters": {
"question": {"ilike": "user question"},
"feedback_good": {"eq": true}
},
"pagination": {"limit": 10}
}'

SDK

from nuclia import sdk
from nuclia_models.events.activity_logs import ActivityLogsAskQuery, EventType, Pagination

kb = sdk.NucliaKB()
query = ActivityLogsAskQuery(
year_month="2024-10",
show=["id", "date", "question", "answer"],
filters={
"question": {"ilike": "user question"},
"feedback_good": {"eq": True}
},
pagination=Pagination(limit=10)
)
kb.logs.query(type=EventType.ASK, query=query)

Filtering by list values

Use isin or isnotin:

filters={"answer": {"isin": ["alpha", "gamma"]}}

Filtering by audit_metadata

audit_metadata is a customizable dictionary. Use the key operator to target specific keys:

query = ActivityLogsAskQuery(
year_month="2024-10",
show=["audit_metadata.environment"],
filters={
"audit_metadata": [{"key": "environment", "eq": "prod"}]
},
pagination=Pagination(limit=10)
)

Download

CLI

# Wait for the download URL to be generated (blocking)
nuclia kb logs download --wait --type=ASK --format=NDJSON --query='{
"year_month": "2024-10",
"show": ["id", "date", "question", "answer", "feedback_good"],
"filters": {"question": {"ilike": "user question"}}
}'

# Request download and get notified via email
nuclia kb logs download --type=ASK --format=NDJSON --query='{
"year_month": "2024-10",
"show": ["id", "date", "question", "answer"],
"notify_via_email": true,
"email_address": "address@foo.com"
}'

# Poll for status manually
nuclia kb logs download_status <request_id>

SDK

from nuclia import sdk
from nuclia_models.events.activity_logs import (
DownloadActivityLogsAskQuery, DownloadFormat, EventType,
)

kb = sdk.NucliaKB()
query = DownloadActivityLogsAskQuery(
year_month="2024-10",
show=["id", "date", "question", "answer"],
filters={
"question": {"ilike": "user question"},
"feedback_good": {"eq": True}
},
)
request = kb.logs.download(
type=EventType.ASK, query=query, download_format=DownloadFormat.NDJSON, wait=True
)
print(request.download_url)

REMi

The REMi module monitors the quality of your RAG pipeline. Use it to query logs by REMi scores and track score evolution over time.

Query

Retrieve ask activity logs matching REMi score criteria.

CLI

nuclia kb remi query --query='{
"month": "2024-11",
"context_relevance": {
"value": 0,
"operation": "gt",
"aggregation": "average"
}
}'

SDK

from nuclia import sdk
from nuclia_models.events.remi import RemiQuery, ContextRelevanceQuery

kb = sdk.NucliaKB()
kb.remi.query(
query=RemiQuery(
month="2024-11",
context_relevance=ContextRelevanceQuery(
value=0, operation="gt", aggregation="average"
),
)
)

Optional filters: feedback_good (bool) and status (NO_CONTEXT, ERROR, SUCCESS):

from nuclia_models.events.remi import RemiQuery, ContextRelevanceQuery, Status

kb.remi.query(
query=RemiQuery(
month="2024-11",
context_relevance=ContextRelevanceQuery(value=0, operation="gt", aggregation="average"),
feedback_good=True,
status=Status.SUCCESS,
)
)

Get Event

Fetch full context and score details for a specific event (from a previous query result):

nuclia kb remi get_event --event_id=16987522
kb.remi.get_event(event_id=16987522)

Get Scores

Retrieve REMi score progression over time, aggregated by day, week, or month:

nuclia kb remi get_scores --starting_at=2024-05-01 --to=None --aggregation=day
from nuclia_models.common.utils import Aggregation
from datetime import datetime

output = kb.remi.get_scores(
starting_at=datetime(year=2024, month=5, day=1),
to=None,
aggregation=Aggregation.DAY,
)