Langflow Components¶
Flexible GraphRAG provides 12 custom Langflow components under the Flexible GraphRAG node category. They are thin nodes over the real backend machinery — the ingest/* pipeline functions (run_chunk_pipeline, update_vector / update_search / update_pg_graph / update_rdf_graph), retriever_setup, and the full adapter/store layer (both LangChain and LlamaIndex). Nothing about the pipeline is reimplemented in the components; they marshal inputs and call the same code the FastAPI backend does.
For setup and app-driven flow mode, see Langflow Integration. This page is the developer reference for the components themselves.
Source: flexible-graphrag/langflow_components/flexible_graphrag/.
Architecture¶
- Driven by
.env. Each component reads the backendSettings(). UI fields on a node override the corresponding setting as an environment variable before the system is built, so a flow can tweak per-node behavior without editing.env. - One shared system per run. The system class is always
HybridSearchSystem(the full LangChain + LlamaIndex adapter layer). The Data Source node builds it and stashes it in a process-level run cache; only a short run-key string travels in theDatapayloads along the edges (the live system is not serializable). Downstream nodes look the system back up by run-key. - Ingestion threads the shared system along the edges; query nodes are independent (Hybrid Search and AI Query each build their own system and reconnect to the persistent stores — they are not wired to the ingestion graph).
- Imports. Components import the shared helper by absolute path —
from langflow_components.flexible_graphrag._fg_shared import ...— because Langflow's component loader cannot resolve a bare local import.
The 12 components¶
| # | Node (display name) | name / type |
Icon | In → Out | Role |
|---|---|---|---|---|---|
| 1 | Flexible: Data Source | FlexibleDataSource |
FolderOpen | Trigger (Message) → Data | Entry point: builds the shared system (.env + UI overrides) and resolves a source to file paths. |
| 2 | Flexible: Document Processor | FlexibleDocProcessor |
FileText | Source (Data) → Data | Parse source files into documents via the backend DocumentProcessor (Docling / LlamaParse per .env). |
| 3 | Flexible: Text Splitter | FlexibleTextSplitter |
Scissors | Documents (Data) → Data | Chunk documents (size/overlap and chunker backend from .env). |
| 4 | Flexible: Vector Store | FlexibleVectorStore |
Database | Chunks (Data) → Data | Index chunks into the vector DB (Qdrant / Neo4j / Elasticsearch / … per .env). |
| 5 | Flexible: KG Extraction | FlexibleKGExtraction |
Sparkles | Chunks (Data) → Data | Extract entities/relations once; feeds Property Graph + RDF in parallel. |
| 6 | Flexible: Property Graph | FlexibleGraphStore |
Workflow | KG (Data) → Data | Store the extracted KG into the property-graph DB (Neo4j / … per .env). |
| 7 | Flexible: Search Store | FlexibleSearchStore |
Search | Chunks (Data) → Data | Build the fulltext/BM25 index (BM25 / Elasticsearch / OpenSearch per .env). |
| 8 | Flexible: RDF Graph | FlexibleRDFStore |
Share2 | KG (Data) → Data | Write RDF triples (Fuseki / Oxigraph / GraphDB / Neptune per .env). |
| 9 | Flexible: Ingestion Summary | FlexibleIngestSummary |
ListChecks | Vector/Search/PG/RDF (Data) → Message | Join the store nodes; print a one-shot ingestion summary (feeds ChatOutput for the Playground). |
| 10 | Flexible: Hybrid Search | FlexibleHybridSearch |
Zap | Query (Message) → Data | Hybrid search over all configured modalities (vector, property graph, fulltext, RDF). Standalone. |
| 11 | Flexible: AI Query | FlexibleAIQuery |
MessageSquare | Question (Message) → Data | Answer a question via RAG over the ingested stores. Independent of Hybrid Search. |
| 12 | Flexible: Query Summary | FlexibleQuerySummary |
ListChecks | AI Answer + Search Results (Data) → Message | Combine the AI Query answer + Hybrid Search results into one chat message. |
Ingestion flow topology¶
[ChatInput]
│ (input_value: {source_type, source_config, config_path, …})
▼
Data Source ──► Document Processor ──► Text Splitter ──┬──► Vector Store ─────────────┐
│ │
├──► Search Store ─────────────┤
│ │
└──► KG Extraction ──┬──► Property Graph ─┤
│ │
└──► RDF Graph ──────┤
▼
Ingestion Summary ──► [ChatOutput]
- KG Extraction runs once and fans out to both Property Graph and RDF Graph, so entities/relations aren't extracted twice.
- Any store node whose backend is disabled in
.envis a no-op — the flow still runs end-to-end. - The Ingestion Summary joins whichever store outputs are present (all four are optional inputs) and emits the human-readable summary.
Query flow topology¶
Hybrid Search and AI Query are independent — each builds its own system and reconnects to the persistent stores. There are three query-side flows:
Combined (Playground only — fg_query_flow.json):
[ChatInput: question] ──┬──► Hybrid Search ──► results ─┐
│ ├──► Query Summary ──► [ChatOutput]
└──► AI Query ──────► answer ────┘
App search (fg_search_flow.json): [ChatInput] ──► Hybrid Search
App AI query (fg_aiquery_flow.json): [ChatInput] ──► AI Query
The app uses the dedicated single-branch flows because Langflow runs every branch present in a flow — if the app ran the combined flow for a search it would also fire an AI Query LLM call (and vice-versa). The backend requests the target component's structured output via output_component, so run_search_flow / run_aiquery_flow return the same shapes as direct mode. The combined flow (with Query Summary merging both into one chat message) is kept for the Langflow Playground.
Each query node reuses a warm cached system (run_with_query_system in _fg_shared.py) instead of rebuilding it every request; the cache is invalidated on each ingest (start_run) so queries always reflect the latest data.
Component loading & Langflow 1.10.2 constraints¶
As of v0.7.1 the components load via a Langflow Extension bundle — langflow_components/extension.json plus a [langflow.extensions] entry point in pyproject.toml (flexible-graphrag = "langflow_components"). Langflow discovers it at startup by walking importlib.metadata.distributions(), so installing flexible-graphrag in Langflow's venv is all it takes — no LANGFLOW_COMPONENTS_PATH. The bundle's path (flexible_graphrag) names the sidebar category; "lfx": {"compat": ["1"]} targets Langflow major version 1 (not pinned to one release). Setting LANGFLOW_COMPONENTS_PATH as well is unnecessary and makes Langflow log a bundle-shadowed warning (the installed extension outranks the inline path).
Langflow 1.10.2 requires the langchain shadow fix. This repo's own package is named langchain and wins import langchain in editable/source layouts (its subpackages import as langchain.graph/langchain.llm/…). Langflow 1.10.2's core startup does an uncaught from langchain.agents import create_agent; with our package shadowing the real one, that aborted boot (No module named 'langchain.agents'). langchain/__init__.py runs pkgutil.extend_path so our subpackages resolve first while the real langchain's agents/chains/tools fall through. Without it, Langflow 1.10.2 won't start in an editable/source install (the PyPI wheel install is unaffected — the two langchain dirs merge in site-packages).
Langflow's loader (lfx.custom.validate.prepare_global_scope) only executes top-level import / class / def AST nodes in a component file. Three patterns are required for the components to register on Langflow 1.10.2:
- Secret fields use
SecretStrInput(...), notMessageTextInput(..., password=True)— API keys must be secret inputs. - Imports must be flat and top-level.
try/exceptimport blocks are silently dropped (they causedNameError: Component is not defined). Langflow auto-falls-backlangflow.* → lfx.*. - Input ports are declared with
HandleInput(..., input_types=["Data"])— declaring an input asOutput(...)insideinputs=[...]does not work.
Other loader details:
- Package layout drives discovery. Each component is its own file, and an
__init__.pymust exist at every level (langflow_components/andlangflow_components/flexible_graphrag/) — a missing__init__.pymakes components silently not appear. The subfolder name (flexible_graphrag) becomes the sidebar category. - Icons are Lucide icon names (
icon = "FileText"), and an optionalpriority: intorders the components within the palette (lower first).
Fast iteration without a restart¶
To debug a single component without the regen + Langflow-restart cycle, use Langflow's "New Custom Component" button at the bottom of the left sidebar and paste the component's Python directly — it registers immediately. (The whole package registers permanently via the extension bundle once flexible-graphrag is installed in Langflow's venv.)
Mixing with stock Langflow components¶
The Flexible: * nodes speak LlamaIndex Data/documents; stock Langflow nodes are LangChain-native. If you wire a Flexible component directly to a built-in Langflow node, the document shape differs — LlamaIndex Document(text=…, metadata=…) vs LangChain Document(page_content=…, metadata=…). Convert at the boundary (page_content ↔ text, pass metadata through). The backend already provides these converters (langchain/utils.py) for the ingest pipeline.
Regenerating the bundled flow JSONs¶
Regeneration is only needed when you change internal flexible-graphrag code — the flow JSONs embed the component code, so an edit to a component (or the pipeline it calls) isn't picked up in flow mode until the flows are regenerated. Normal use never needs it: the app drives the existing flows over the Langflow REST API, passing per-run config as the flow's input_value.
The four flows in flows/ are generated from the real registered component templates (via lfx.custom.utils.build_custom_component_template), then assembled into nodes + edges:
fg_ingestion_flow.json— the full ingest pipeline.fg_search_flow.json— ChatInput → Hybrid Search (the app's search).fg_aiquery_flow.json— ChatInput → AI Query (the app's AI query).fg_query_flow.json— combined Hybrid Search + AI Query → Query Summary (Langflow Playground only; the app doesn't run it).
The dedicated search / AI-query flows exist because Langflow runs every branch present in a flow — running the combined flow for a search would also fire an AI Query LLM call. run_search_flow / run_aiquery_flow target the dedicated flows (falling back to the combined one if a file is missing).
Regenerate all four from the live components:
It validates each with lfx.graph.graph.base.Graph.from_payload(data).prepare() before writing. In app-driven flow mode the backend deletes and re-uploads the flow matching the JSON's name on every startup (see Langflow Integration), so regenerate rather than hand-editing the running flow in the UI. Re-run the generator whenever you edit a component (its code is embedded in the flow JSON), then restart the backend.
Creating a Langflow API key programmatically¶
When Langflow requires auth (LANGFLOW_AUTO_LOGIN=false), the backend needs a key in LANGFLOW_API_KEY. The UI route is Settings → Langflow API Keys → Add New (see Langflow Integration → Authentication); to mint one from code, call the Langflow API:
import httpx
LF = "http://localhost:7860"
# 1) Get a session token.
# AUTO_LOGIN on: GET /api/v1/auto_login
# AUTO_LOGIN off: POST /api/v1/login (form: username/password = your superuser creds)
tok = httpx.get(f"{LF}/api/v1/auto_login").json()["access_token"]
# auth mode instead:
# tok = httpx.post(f"{LF}/api/v1/login", data={"username": "admin", "password": "..."}).json()["access_token"]
# 2) Create the key — POST /api/v1/api_key/ ; the response's "api_key" is the FULL key (returned once).
r = httpx.post(f"{LF}/api/v1/api_key/",
headers={"Authorization": f"Bearer {tok}"},
json={"name": "flexible-graphrag"})
print(r.json()["api_key"]) # → set this as the backend's LANGFLOW_API_KEY
GET /api/v1/api_key/ lists keys (masked); DELETE /api/v1/api_key/{id} removes one. Keys live in langflow.db — persist it (the langflow_data volume in Docker) if they should survive restarts. The langflow api-key create CLI is the terminal equivalent.
File map¶
| Path | Contents |
|---|---|
langflow_components/flexible_graphrag/*.py |
The 12 component classes (one per file) + __init__.py exports. |
langflow_components/flexible_graphrag/_fg_shared.py |
Run cache, system builder (build_system), the warm query-system cache (get_query_system / run_with_query_system, cleared on ingest by start_run), has_ingest_chunks / ingest_chunk_count helpers. |
langflow_components/generate_flows.py |
Regenerates all four flow JSONs from the live component classes (run in the Langflow venv). |
flows/fg_ingestion_flow.json, fg_search_flow.json, fg_aiquery_flow.json, fg_query_flow.json |
The bundled flows the app uploads (ingest, search, AI query, + combined Playground query). |
flow_service.py (backend) |
FlowService — drives the flows over the Langflow REST API in app-driven mode (uses the dedicated search / AI-query flows). |