> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/security-testing-orchestration/use-sto/sto-scanner-configuration/semgrep/semgrep-scanner-reference.md).

# Semgrep step configuration

You can scan your code repositories using [Semgrep](https://www.semgrep.com) and ingest the results into STO.

For a quick introduction, go to the [SAST code scans using Semgrep](/security-testing-orchestration/use-sto/sto-scanner-configuration/semgrep/sast-scan-semgrep.md) tutorial.

### Important notes for running Semgrep scans in STO <a href="#important-notes-for-running-semgrep-scans-in-sto" id="important-notes-for-running-semgrep-scans-in-sto"></a>

* This integration uses the [Semgrep Engine](https://github.com/semgrep/semgrep), which is open-source and licensed under [LGPL 2.1](https://tldrlegal.com/license/gnu-lesser-general-public-license-v2.1-\(lgpl-2.1\)).

  To run scans using a licensed version of [Semgrep Code](https://semgrep.dev/products/semgrep-code), add your Semgrep token in the [Access token](#access-token) field.
* STO Semgrep steps include the following rulesets by default:

  * [auto](https://semgrep.dev/p/auto)
  * [bandit](https://semgrep.dev/p/bandit)
  * [brakeman](https://semgrep.dev/p/brakeman)
  * [eslint](https://semgrep.dev/p/eslint)
  * [findsecbugs](https://semgrep.dev/p/findsecbugs)
  * [flawfinder](https://semgrep.dev/p/flawfinder)
  * [gosec](https://semgrep.dev/p/gosec)
  * [phps-security-audit](https://semgrep.dev/p/phpcs-security-audit)
  * [security-code-scan](https://semgrep.dev/p/security-code-scan)

  Some rulesets include Pro rules that are available only with a paid version of Semgrep. For more information, go to the [Semgrep Registry](https://semgrep.dev/explore).
* If you want to add trusted certificates to your scan images at runtime, you need to run the scan step with root access.

  You can set up your STO scan images and pipelines to run scans as non-root and establish trust for your proxies using custom certificates. For more information, go to [Configure your pipeline to use STO images from private registry](/security-testing-orchestration/troubleshooting-and-resources/sto-use-cases/set-up-sto-pipelines/configure-pipeline-to-use-sto-images-from-private-registry.md).
* The following topics contain useful information for setting up scanner integrations in STO:
  * [What's supported in STO](/security-testing-orchestration/new-to-sto/sto-whats-supported/sto-deployments.md)
  * [Security Testing Orchestration FAQs](/security-testing-orchestration/troubleshooting-and-resources/faqs.md)
  * [Optimize STO pipelines](/security-testing-orchestration/troubleshooting-and-resources/sto-use-cases/set-up-sto-pipelines/optimize-sto-pipelines.md)

### Set-up workflows <a href="#set-up-workflows" id="set-up-workflows"></a>

<details>

<summary>Orchestration scans</summary>

To scan a code repository, you need [Harness Code Repository](https://app.gitbook.com/s/oTP6ysjQCQlt5GlFHRfs/README) or a [Harness connector](/harness-ai/use-harness-platform/connectors/code-repositories/ref-source-repo-provider.md) to your Git service.

**Add the Semgrep scanner**

Do the following:

1. Add a **Build** or **Security** stage to your pipeline.
2. Configure the stage to point to the [codebase](/continuous-integration/use-harness-ci/use-harness-ci/codebase-configuration/create-and-configure-a-codebase.md) you want to scan.
3. Add a Semgrep step to the stage.

**Set up the Semgrep scanner**

**Required settings**

1. Scan mode = [Orchestration](#scan-mode)
2. Target and Variant Detection = [Auto](#detect-target-and-variant)

**Optional settings**

* [Fail on Severity](#fail-on-severity) — Stop the pipeline if the scan detects any issues at a specified severity or higher
* [Log Level](#log-level) — Useful for debugging

**Scan the repository**

Save your pipeline and then select **Run**.

The pipeline scans your code repository and then shows the results in [Vulnerabilities tab](/security-testing-orchestration/use-sto/sto-security-issues/view-scan-results.md).

</details>

<details>

<summary>Ingestion scans</summary>

**Add a shared path for your scan results**

1. Add a **Build** or **Security** stage to your pipeline.
2. In the stage **Overview**, add a shared path such as `/shared/scan_results`.

**Copy scan results to the shared path**

There are two primary workflows to do this:

* Add a Run step that runs a Semgrep scan from the command line and then copies the results to the shared path.
* Copy results from a Semgrep scan that ran outside the pipeline.

For more information and examples, go to [Ingestion scans](/security-testing-orchestration/new-to-sto/key-concepts/ingest-scan-results-into-an-sto-pipeline.md).

**Set up the Semgrep scanner**

Add a Semgrep step to the stage and set it up as follows.

**Required settings**

1. [Scan mode](#scan-mode) = Ingestion
2. [Target name](#name) — Usually the repo name
3. [Target variant](#name) — Usually the scanned branch. You can also use a [runtime input](/harness-ai/use-harness-platform/variables-and-expressions/runtime-input-usage.md) and specify the branch at runtime.
4. [Ingestion file](#ingestion-file) — For example, `/shared/scan_results/semgrep-scan.json`

**Optional settings**

* [Fail on Severity](#fail-on-severity) — Stop the pipeline if the scan detects any issues at a specified severity or higher
* [Log Level](#log-level) — Useful for debugging

**Scan the repository**

Save your pipeline and then select **Run**.

The pipeline scans your code repository and then shows the results in [Vulnerabilities tab](/security-testing-orchestration/use-sto/sto-security-issues/view-scan-results.md).

</details>

### Semgrep step configuration <a href="#semgrep-step-configuration" id="semgrep-step-configuration"></a>

The recommended workflow is to add a Semgrep step to a Security Tests or CI Build stage and then configure it as described below.

#### Scan <a href="#scan" id="scan"></a>

**Scan Mode**

* **Orchestration** Configure the step to [run a scan](/security-testing-orchestration/new-to-sto/key-concepts/run-an-orchestrated-scan-in-sto.md) and then ingest, normalize, and deduplicate the results.
* **Ingestion** Configure the step to [read scan results from a data file](/security-testing-orchestration/new-to-sto/key-concepts/ingest-scan-results-into-an-sto-pipeline.md) and then ingest, normalize, and deduplicate the data.

**Scan Configuration**

You can use this setting to select the set of Semgrep rulesets to include in your scan:

* **Default** Include the following rulesets:
  * [bandit](https://semgrep.dev/p/bandit)
  * [brakeman](https://semgrep.dev/p/brakeman)
  * [eslint](https://semgrep.dev/p/eslint)
  * [findsecbugs](https://semgrep.dev/p/findsecbugs)
  * [flawfinder](https://semgrep.dev/p/flawfinder)
  * [gosec](https://semgrep.dev/p/gosec)
  * [phps-security-audit](https://semgrep.dev/p/phpcs-security-audit)
  * [security-code-scan](https://semgrep.dev/p/security-code-scan)
* **No default CLI flags** Run the `semgrep` scanner with no additional CLI flags. This setting is useful if you want to specify a custom set of rulesets in **Additional CLI flags**.
* **p/default** Run the scan with the [default ruleset](https://semgrep.dev/p/default) configured for the Semgrep scanner.
* **Auto only** Run the scan with the [recommended rulesets specific to your project](https://semgrep.dev/p/auto).
* **Auto and Ported security tools** Include the following rulesets:
  * [auto](https://semgrep.dev/p/auto)
  * [brakeman](https://semgrep.dev/p/brakeman)
  * [eslint](https://semgrep.dev/p/eslint)
  * [findsecbugs](https://semgrep.dev/p/findsecbugs)
  * [flawfinder](https://semgrep.dev/p/flawfinder)
  * [gitleaks](https://semgrep.dev/p/gitleaks)
  * [gosec](https://semgrep.dev/p/gosec)
  * [phps-security-audit](https://semgrep.dev/p/phpcs-security-audit)
  * [security-code-scan](https://semgrep.dev/p/security-code-scan)
* **Auto and Ported security tools except p/gitleaks**

#### Target <a href="#target" id="target"></a>

**Type**

* **Repository** Scan a codebase repo.

  In most cases, you specify the codebase using a code repo connector that connects to the Git account or repository where your code is stored. For information, go to [Configure codebase](/continuous-integration/use-harness-ci/use-harness-ci/codebase-configuration/create-and-configure-a-codebase.md).

**Target and variant detection**

When **Auto** is enabled for code repositories, the step detects these values using `git`:

* To detect the target, the step runs `git config --get remote.origin.url`.
* To detect the variant, the step runs `git rev-parse --abbrev-ref HEAD`. The default assumption is that the `HEAD` branch is the one you want to scan.

Note the following:

* **Auto** is not available when the **Scan Mode** is **Ingestion**.
* By default, **Auto** is selected when you add the step. You can change this setting if needed.

**Name**

The identifier for the [target](/security-testing-orchestration/new-to-sto/key-concepts/targets-and-baselines.md), such as `codebaseAlpha` or `jsmith/myalphaservice`. Descriptive target names make it much easier to navigate your scan data in the STO UI.

It is good practice to [specify a baseline](/security-testing-orchestration/new-to-sto/key-concepts/targets-and-baselines.md#every-target-needs-a-baseline) for every target.

**Variant**

The identifier for the specific variant to scan. This is usually the branch name, image tag, or product version. Harness maintains a historical trend for each variant.

**Workspace**

The workspace path on the pod running the scan step. The workspace path is `/harness` by default.

You can override this if you want to scan only a subset of the workspace. For example, suppose the pipeline publishes artifacts to a subfolder `/tmp/artifacts` and you want to scan these artifacts only. In this case, you can specify the workspace path as `/harness/tmp/artifacts`.

Additionally, you can specify individual files to scan as well. For instance, if you only want to scan a specific file like `/tmp/iac/infra.tf`, you can specify the workspace path as `/harness/tmp/iac/infra.tf`

#### Ingestion File <a href="#ingestion-file" id="ingestion-file"></a>

The path to your scan results when running an [Ingestion scan](/security-testing-orchestration/new-to-sto/key-concepts/ingest-scan-results-into-an-sto-pipeline.md), for example `/shared/scan_results/myscan.latest.sarif`.

* The data file must be in a [supported format](/security-testing-orchestration/new-to-sto/sto-whats-supported/scanners.md#supported-ingestion-formats) for the scanner.
* The data file must be accessible to the scan step. It's good practice to save your results files to a [shared path](/continuous-integration/new-to-harness-ci/key-concepts.md#stages) in your stage. In the visual editor, go to the stage where you're running the scan. Then go to **Overview** > **Shared Paths**. You can also add the path to the YAML stage definition like this:

  ```yaml
      - stage:
        spec:
          sharedPaths:
            - /shared/scan_results
  ```

#### Access Token <a href="#access-token" id="access-token"></a>

The access token to log in to the scanner. This is usually a password or an API key.

You should create a Harness text secret with your encrypted token and reference the secret using the format `<+secrets.getValue("my-access-token")>`. For more information, go to [Add and Reference Text Secrets](/harness-ai/use-harness-platform/secrets/add-use-text-secrets.md).

#### Log Level <a href="#log-level" id="log-level"></a>

The minimum severity of the messages you want to include in your scan logs. You can specify one of the following:

* **DEBUG**
* **INFO**
* **WARNING**
* **ERROR**

#### Additional CLI flags <a href="#additional-cli-flags" id="additional-cli-flags"></a>

Use this field to run the [`semgrep`](https://semgrep.dev/docs/cli-reference/) scanner with flags such as:

`--severity=ERROR --use-git-ignore`

With these flags, `semgrep` considers only ERROR severity rules and ignores files included in `.gitignore`.

{% hint style="warning" %}
Passing additional CLI flags is an advanced feature. Harness recommends the following best practices:

* Test your flags and arguments thoroughly before you use them in your Harness pipelines. Some flags might not work in the context of STO.
* Don't add flags that are already used in the default configuration of the scan step.

  To check the default configuration, go to a pipeline execution where the scan step ran with no additional flags. Check the log output for the scan step. You should see a line like this:

  `Command [ scancmd -f json -o /tmp/output.json ]`

  In this case, don't add `-f` or `-o` to **Additional CLI flags**.
  {% endhint %}

#### Fail on Severity <a href="#fail-on-severity" id="fail-on-severity"></a>

Every STO scan step has a **Fail on Severity** setting. If the scan finds any vulnerability with the specified [severity level](/security-testing-orchestration/new-to-sto/key-concepts/severities.md) or higher, the pipeline fails automatically. You can specify one of the following:

* **`CRITICAL`**
* **`HIGH`**
* **`MEDIUM`**
* **`LOW`**
* **`INFO`**
* **`NONE`** — Do not fail on severity

The YAML definition looks like this: `fail_on_severity : critical # | high | medium | low | info | none`

#### Settings <a href="#settings" id="settings"></a>

You can use this field to specify environment variables for your scanner.

#### Additional Configuration <a href="#additional-configuration" id="additional-configuration"></a>

The fields under **Additional Configuration** vary based on the type of infrastructure. Depending on the infrastructure type selected, some fields may or may not appear in your settings. Below are the details for each field

* Override Security Test Image
  * [Container Registry](/security-testing-orchestration/troubleshooting-and-resources/sto-use-cases/set-up-sto-pipelines/configure-pipeline-to-use-sto-images-from-private-registry.md#step-level-override)
  * [Image Tag](/security-testing-orchestration/troubleshooting-and-resources/sto-use-cases/set-up-sto-pipelines/configure-pipeline-to-use-sto-images-from-private-registry.md#step-level-override)
* [Privileged](/continuous-integration/use-harness-ci/use-harness-ci/manage-dependencies/background-step-settings.md#privileged)
* [Image Pull Policy](/continuous-integration/use-harness-ci/use-harness-ci/manage-dependencies/background-step-settings.md#image-pull-policy)
* [Run as User](/continuous-integration/use-harness-ci/use-harness-ci/manage-dependencies/background-step-settings.md#run-as-user)
* [Set Container Resources](/continuous-integration/use-harness-ci/use-harness-ci/manage-dependencies/background-step-settings.md#set-container-resources)
* [Timeout](/continuous-integration/use-harness-ci/use-harness-ci/run-step-settings.md#timeout)

#### Advanced settings <a href="#advanced-settings" id="advanced-settings"></a>

In the **Advanced** settings, you can use the following options:

* [Conditional Execution](/harness-ai/use-harness-platform/pipelines/step-skip-condition-settings.md)
* [Failure Strategy](/harness-ai/use-harness-platform/pipelines/failure-handling/define-a-failure-strategy-on-stages-and-steps.md)
* [Looping Strategy](/harness-ai/use-harness-platform/pipelines/looping-strategies/looping-strategies-matrix-repeat-and-parallelism.md)
* [Policy Enforcement](/harness-ai/use-harness-platform/governance/policy-as-code/harness-governance-overview.md)

### Configure Semgrep as a Built-in Scanner <a href="#configure-semgrep-as-a-built-in-scanner" id="configure-semgrep-as-a-built-in-scanner"></a>

The Semgrep scanner is available as a [built-in scanner](/security-testing-orchestration/use-sto/set-up-sto-scans/built-in-scanners.md) in STO. Configuring it as a built-in scanner enables the step to automatically perform scans using the free version without requiring any licenses. Follow these steps to set it up:

1. Search for **SAST** in the step palette or navigate to the **Built-in Scanners** section and select the **SAST** step.
2. Expand the **Additional CLI Flags** section if you want to configure optional CLI flags.
3. Click **Add Scanner** to save the configuration.

The scanner will automatically use the free version, detect scan targets, and can be further configured by clicking on the step whenever needed.

### Proxy settings <a href="#proxy-settings" id="proxy-settings"></a>

This step supports private network connectivity if you're using Harness Cloud infrastructure. For information on connectivity options, see [Private network connectivity options](/harness-ai/use-harness-platform/references/private-network-connectivity/private-network-connectivity.md). When using proxy configurations, the `HTTPS_PROXY` and `HTTP_PROXY` variables are automatically set to route traffic through the secure tunnel. If there are specific addresses that you want to bypass the proxy, you can define those in the `NO_PROXY` variable. This can be configured in the **Settings** of your step.

If you need to configure a different proxy, you can manually set the `HTTPS_PROXY`, `HTTP_PROXY`, and `NO_PROXY` variables in the **Settings** of your step.

**Definitions of Proxy variables:**

* `HTTPS_PROXY`: Specify the proxy server for HTTPS requests, example `https://sc.internal.harness.io:30000`
* `HTTP_PROXY`: Specify the proxy server for HTTP requests, example `http://sc.internal.harness.io:30000`
* `NO_PROXY`: Specify the domains as comma-separated values that should bypass the proxy. This allows you to exclude certain traffic from being routed through the proxy.

### YAML pipeline example <a href="#yaml-pipeline-example" id="yaml-pipeline-example"></a>

The following pipeline example illustrates an orchestration workflow. It consists of a Semgrep step that scans a code repository and then ingests, normalizes, and deduplicates the results.

![](/files/cfU1fZWwszPU4mxqKjgg)

```yaml
pipeline:
  name: semgrep-orch-test
  identifier: semgreporchtest
  projectIdentifier: default
  orgIdentifier: default
  tags: {}
  properties:
    ci:
      codebase:
        connectorRef: YOUR_GIT_CONNECTOR_ID
        repoName: YOUR_GIT_REPO_NAME
        build: <+input>
  stages:
    - stage:
        name: semgrep-orch
        identifier: semgreporch
        description: ""
        type: SecurityTests
        spec:
          cloneCodebase: true
          platform:
            os: Linux
            arch: Amd64
          runtime:
            type: Cloud
            spec: {}
          execution:
            steps:
              - step:
                  type: Semgrep
                  name: Semgrep_1
                  identifier: Semgrep_1
                  spec:
                    mode: orchestration
                    config: default
                    target:
                      type: repository
                      detection: auto
                    advanced:
                      log:
                        level: info
```
